Why I retired shelpa

shelpa was an MCP (Model Context Protocol) virtual pipeline shell that restricted LLM-agent writes to tee. It mirrored all writes to .shelpa/ to preserve edit history.

I implemented path-escape checks and memory limits, but models did not consistently use the tool, so I retired it. LLM integration required checking actual tool choices as well as implementing restrictions.

Routing writes through tee

Separating reads from writes

The design replaced write_file and edit_file with UNIX pipeline tee as the only write mechanism:

  rg "pattern" src/ | awk '{print $2}' | tee output.txt
  
  • rg, awk, sed, jq, etc. are used as read buffers
  • The final write goes through tee — the only write mechanism
  • tee writes are automatically mirrored to .shelpa/, preserving edit history

A hypothesis about learned shell knowledge

I made the interface resemble a shell so models could reuse learned command knowledge.

Tool names followed shell vocabulary such as tail, rg, awk, sed, and tee:

Tool NameRole
shelpa_pipeExecute pipeline commands
shelpa_teeFile writing (via tee)
shelpa_writeDirect writing (tee wrapper)
shelpaNavigation (cd, pwd)

The help output was:

  shelpa-mcp (MCP stdio server)
Usage:
  shelpa-mcp [--root <ROOT>] [--help]
Notes:
  - This binary speaks MCP over stdio. It does not serve HTTP.
  - Use your MCP client to call tools below.
  - --root sets the workspace root directory for all tool calls.
  - cwd defaults to the workspace root if not provided.
Tools:
  - shelpa_pipe    Execute a virtual safety pipeline
  (tail  rg  awk  sed  tr  jq  wc  tee  fd  ls  sort  head)
  - shelpa_write   Execute a virtual safety tee (auto guard, save override history)
Allowed pipeline commands: tail rg awk sed tr jq wc tee fd
Navigation commands (single-stage only): pwd  cd <path>  ls [path]
Pipes only. No redirects (> >>). No sed -i. No awk file output. Save via tee only.
CRITICAL: Never use standard file editing tools (such as write_file, replace, etc.)
  - always use the specified tool exclusively.
  

I expected models to recognize shelpa_pipe as a pipeline tool. The help’s CRITICAL line also instructed them to use shelpa instead of standard editing tools.

Security Design

Vulnerability 1: Path Escape via cwd Option

tool_pipe canonicalized cwd without checking whether it stayed inside the workspace root.

  // After fix
let canonical = fs::canonicalize(&resolved)?;
if !canonical.starts_with(root.canonicalize()?) {
    return Err(GuardViolation::new(
        GuardReason::PathEscape,
        "cwd must be within workspace root"
    ));
}
  

Vulnerability 2: Path Escape Within .shelpa

shelpa_write_path joined the target directly, allowing ../ to escape .shelpa/.

  // After fix: compute relative path from canonicalized real path
pub fn shelpa_write_path(workspace_root: &Path, real_path: &Path) -> Result<PathBuf, GuardViolation> {
    let real_canonical = fs::canonicalize(real_path)?;
    let relative = real_canonical.strip_prefix(workspace_root.canonicalize()?)
        .map_err(|_| GuardViolation::new(GuardReason::PathEscape, "Path is outside workspace"))?;
    Ok(workspace_root.join(".shelpa").join(relative))
}
  

Vulnerability 3: Unbounded Memory Buffering in tee

tee buffered all output in memory, which could cause OOM on large output. I added two limits:

  1. Hard limit (1MB): Excess is truncated, truncated=true set in metadata
  2. Soft limit (2KB, approval gate): Returns APPROVAL_REQUIRED error when exceeded, requiring explicit user confirmation
  if out_data.len() > TEE_APPROVAL_BYTES && !confirm_oversize {
    return Err(GuardViolation::new(
        GuardReason::ApprovalRequired,
        format!("tee would write {} bytes, exceeding the {}-byte threshold. \
                 Re-invoke with confirm_oversize=true to proceed.",
                out_data.len(), TEE_APPROVAL_BYTES)
    ));
}
  

Audit Mirror Design

tee mirrored the real file content to .shelpa/ and added timestamp separators on overwrite:

  --- shelpa:overwrite ts=2026-02-25T13:17:30Z record_id=1772025450395572000 ---
(new content)
  

Operational problems

Models did not consistently use the tool

Models did not keep following the instruction to write through shelpa.

Specifically:

  • Low tool selection priority: LLMs default to using write_file and edit_file. Even when shelpa usage was enforced via system prompts, models would forget as context grew longer
  • Pipeline syntax construction errors: Frequent failures in correctly building pipelines like rg pattern | awk '{print $2}' | tee output.txt. Quote escaping was a common point of failure
  • Shell disguise didn’t fool the model: The intent was to trick LLMs with shell-style naming like shelpa_pipe and shelpa_tee, but models didn’t recognize them as “a type of shell.” The scheme to leverage pre-trained shell knowledge had less effect than expected

Findings from this implementation

1. Instructions did not stabilize tool choice

The tested models favored existing tool-use patterns. System prompts alone did not establish this custom interface.

2. Shell-style names were insufficient

Names such as shelpa_pipe and shelpa_tee did not change the preference for write_file. Models sometimes sent one-liners through the pipe, but reviewing them added work.

3. Audit mirrors helped recovery

Mirrors in .shelpa/ helped post-hoc inspection. Pattern matching located checkpoints for recovery after an agent overwrote an important file.