Nine MCP servers for familiar

At the time of this record, familiar used 9 MCP servers, all written in Rust with stdio JSON-RPC 2.0. I built roughly 1-2 per day, retired some, and kept improving the useful ones. These are the tools used to integrate LLMs into development work.

Zed MCP server list -- all 9 servers green
Zed MCP settings -- all servers connected. 41 tools total

Homelab architecture

The homelab running the MCP servers:

Homelab 4-node architecture -- edge, storage, desktop, compute
Homelab architecture. edge (RouterOS), storage (Ubuntu 24/7), desktop (macOS), compute (EPYC + Blackwell x2). MCP servers run on desktop, connecting to compute's LLM backends and storage's data platform
MCP ServerToolsSummary
voracle4Obsidian vault semantic search + web research
argus1Cross-repository code intelligence
ctree8Code tree structure analysis + checkpoints
dgs17secret
memento2secret
okitegami1Discord leave-a-letter
smng2Homelab secret management
sqlimit-15Read-only DB exploration

resolve-inference is integrated into familiar and no longer runs as a separate MCP server. Its design history appears below.


voracle: vault search and web research

Semantic search and web research for Obsidian vaults. ONNX embedding and ColBERT reranking run locally. research collects web findings through Brave Search API and stores them in the vault. Details are in a separate article.

Tools: search, read, research, related


argus: cross-repository symbol search

Indexes all Gitea repositories, extracts symbols with tree-sitter, and reranks results with ColBERT MaxSim.

  CLI (index + embed)           MCP Server (query)
     |                              |
     v                              v
[Gitea API: repos/search]    [CWD auto-index on startup]
     |                              |
     v                              v
[git clone/fetch]             [argus.see tool]
     |                              |
     v                              v
[tree-sitter: symbols]        [adaptive lexical filter -> ColBERT rerank]
     |                              |
     v                              v
[ColBERT embed (batch)]       [Gemini summarize]
     |                              |
     v                              v
[redb: symbols+embeddings]    [hit tracking + NATS pulse]
  

Embedding uses three tiers:

  1. Batch: argus embed pre-computes all symbol embeddings at index time
  2. Hot-hit: symbols accessed 3+ times re-embed at query time to stay fresh
  3. On-demand: uncached symbols embed at query time and persist

Tools: see (specify expected symbols, modules, or specs for cross-repo search)


ctree: structure, dependencies, and revisions

.ctree.toml defines the scope. check analyzes symbols and dependencies, while checkpoint records work milestones.

Tools: check, checkpoint, get_affected, get_depends, get_symbol, get_text, get_revs, get_baseline

In familiar development, checkpoint records each code change and get_affected checks impact before the next implementation. Obsidian work logs contain repeated ctree__check and ctree__get_symbol calls.

Most replies are about 5 tokens or contain symbol names and hashes. Actual code is loaded only when needed through get_text. This limits context use while exposing dependencies. I previously used Serena alongside these tools; it became unnecessary once this setup stabilized. After adding my own Gitea and distribution registry, argus provided references to IaC, architecture, and codebase patterns during coding.

Obsidian work log searching for ctree__check -- agent calls ctree tools in sequence right after thinking
Agent work log. Right after thinking, it calls ctree__get_symbol, ctree__get_affected, ctree__check, ctree__get_depends, ctree__get_text in sequence

dgs – 17 tools (non-public)

The largest tool count among the 9 servers. Design and implementation details are non-public.


memento – 2 tools (non-public)

An MCP for AI agent management. Design and implementation details are non-public.


okitegami: asynchronous Discord updates

Posts LLM-agent progress to Discord at work milestones. It has one tool.

The agent leaves a progress note. I reply when convenient, and the agent reads any reply when it posts its next note.

  1. Leave a letter about current work results on Discord
2. While you're at it, check if past letters have replies
3. If no replies, do nothing
  

Posting a new note triggers the read. There is no polling loop for replies.

okitegami Discord thread -- tegami APP letter with ksh3 reply
A letter thread from tegami. summary, worked, result, concerns, tests are structured. ksh3 replied a few hours later

During development, the agent creates threads in #tegami. I review them later and leave comments. Even a reply a month later is picked up on the next tegami call, when the agent may ask about it or make a correction.

An immediate reply is not required. This connects long autonomous work with human review asynchronously.

Tools: tegami


smng: credentials and connections

Manages homelab credentials, endpoints, and connection details. It retrieves service authentication information within *.home.arpa.

Tools: list_services, get_secret


sqlimit-1: database reads limited to one row

Targets only familiar’s main DB and returns at most one row per query. The tool always wraps the query in a subquery rather than relying on the LLM to write LIMIT 1:

  SELECT * FROM (<user_sql>) AS _q LIMIT 1
  

The dedicated PostgreSQL connection uses default_transaction_read_only = on. The tool also rejects non-SELECT statements for read-only use.

Connection details can come from smng. sqlimit-1 also versions schemas: dump_schema retrieves the schema and creates a new version when it changes. This supports inspection with little context.

smng supplies credentials; sqlimit-1 handles database reads and schema records. Each has a separate responsibility.

Tools: tables, describe, sample, query1, dump_schema


resolve-inference: path resolution for ENOENT errors

Built for cases where an LLM guesses a path and gets ENOENT. It selects paths using a filesystem index, lexical scoring, and ColBERT MaxSim. It is now integrated into familiar.

Query history improved accuracy. A ring buffer holds the last 5 resolved paths and boosts candidates by directory and package affinity:

  Go monorepo (106 files, client.go exists in 4-11 locations):
  With history:    85.0% (17/20)
  Without history: 35.0% (7/20)
  Delta:           +50.0pp
  

I evaluated an ensemble of 3 ColBERT models. All scored 87.3%, while latency increased 3.7x. Changing models did not improve this test, so I removed the ensemble and focused on lexical scoring and query-history correlation.


Retired and integrated tools

Most retired tools were integrated elsewhere. resolve-inference moved into familiar; the early shelpa filesystem ideas were distributed across ctree and dgs.

Rust and stdio JSON-RPC 2.0 let me build a working MCP server in about a day. I try it, retire what is unnecessary, and improve what works.


Shared design

The servers share these principles:

  • Rust + stdio JSON-RPC 2.0: fast process startup, small memory footprint
  • Single responsibility: one tool solves one problem. No feature creep
  • Born from real agent work: built by working backwards from “this is where I got stuck”
  • Graceful degradation: if one tool goes down, others keep running. familiar’s main flow doesn’t stop

41 tools run across 9 processes, using tens of MB in total. Rust’s small binaries and low runtime overhead suit this arrangement of small parallel processes.