Working features and the Go migration

I built a Rust (axum) proxy to expose local LLMs through an OpenAI-compatible API. About 1,600 lines handled OpenAI/Ollama clients, SSE streaming, and model-name-based routing. HTTP, Domain, and Infra were separate layers.

NATS relay, Dagster oneshot jobs, RAG, and Qdrant semantic caching remained stubs. Coordinating SSE, subscriptions, and database writes made the Rust implementation costly, so I moved to Go (Gin).

The Rust types and specifications became the basis for the Go implementation. Goroutines and channels made the same control flow easier to write.

This article is the prequel to the full design record of the Go-based AI orchestration platform.


Prerequisites

  • Language: Rust 2024 edition
  • Framework: axum 0.8, tokio 1.48
  • Serialization: serde / serde_json
  • HTTP client: reqwest 0.12
  • Messaging (design only): async-nats 0.45 (stubbed)
  • Data stores (design only): sqlx 0.8 (PostgreSQL), qdrant-client 1.16 (stubbed)
  • Logging: tracing / tracing-subscriber
  • Containers: Docker (proxy + PostgreSQL 18 + Qdrant)
  • Bind address: 0.0.0.0:8080

Design Decisions

Layered Architecture

We separated HTTP, Domain, and Infrastructure according to Clean Architecture dependency rules.

  Client (CLI / IDE / API consumer)
    |
    v
HTTP Layer [bk/http/]
    |-- routes.rs         endpoint registration
    |-- handlers_chat.rs  OpenAI/Ollama chat handlers
    |-- handlers_embeddings.rs  embeddings handler
    |-- error.rs          OpenAI-compatible error mapping
    |
    v
Domain Layer [bk/domain/]
    |-- chat.rs     ChatCompletionRequest/Response, ChatRoute (Direct/Rag/Workflow)
    |-- rag.rs      RagParams, EmbeddingsRequest/Response
    |-- workflow.rs JobRunId
    |
    v
Service Layer [bk/services/]
    |-- ChatService trait     -> HttpChatService (prod) / StubChatService (dev)
    |-- EmbeddingsService trait -> StubEmbeddingsService
    |-- RagService trait      -> StubRagService
    |
    v
Infrastructure Layer [bk/infra/]
    |-- llm.rs          HttpClient (backend LLM calls)
    |-- qdrant_client.rs QdrantClient (vector search)
    |-- auth.rs         ApiKeyAuthorizer
  

Domain has no external dependencies. Backend and storage changes stay in Infra. The Go version kept these boundaries.

Dual OpenAI/Ollama Compatibility

One entry point serves OpenAI and Ollama clients.

MethodPathFormat
POST/v1/chat/completionsOpenAI compatible
POST/v1/embeddingsOpenAI compatible
GET/v1/modelsOpenAI compatible
POST/api/chatOllama compatible
GET/api/tagsOllama compatible
GET/healthHealth check
GET/readyReadiness

The Ollama handler converts DTOs before calling the shared ChatService.

Multi-Backend Routing

The proxy selects a backend automatically based on model name prefix.

  // HttpChatService
fn pick_client(&self, model: &str) -> &HttpClient {
    // model prefix -> route-specific client, fallback -> default client
}
  

Configured route clients are matched by model name; unmatched requests use the default. Routing stays in Service rather than HTTP.

ChatRoute Branching

Custom headers on the request explicitly select the processing path.

  pub enum ChatRoute {
    Direct,    // forward to backend LLM
    Rag,       // vector search + context injection + LLM
    Workflow,  // delegate to Dagster pipeline
}
  

x-workspace identifies a RAG workspace, x-pipeline a workflow, and x-correlation-id a trace. These headers express direct inference, RAG, and workflows through /v1/chat/completions.

SSE Streaming

The service layer returns one of three response types.

  pub enum ChatServiceResponse {
    Once(ChatCompletionResponse),       // non-stream: single JSON
    Stream(Pin<Box<dyn Stream<...>>>),  // chunked: OpenAI-compatible
    StreamRaw(Pin<Box<dyn Stream<...>>>), // raw text: backend passthrough
}
  

Stream mode sends data: {...}\n\n SSE frames and terminates with [DONE]. StreamRaw passes through the backend SSE response without transformation.


Implementation

What Worked

The following prototype functions worked.

  • /v1/chat/completions — OpenAI-compatible streaming and non-streaming responses
  • /api/chat — Ollama compatibility with internal DTO conversion
  • /v1/models — model listing from backends (both OpenAI and Ollama formats)
  • Multi-backend routing — automatic client selection by model name prefix
  • SSE streaming — parsing SSE lines from reqwest byte streams and relaying them
  • Structured logging — tracing spans per handler recording model name, stream flag, and message count

What Remained Stubbed

These functions had traits and specifications but remained stubs.

FeatureStatusDesign Position
RAG context buildingStubRagService (returns fixed string)Qdrant + PostgreSQL pgvector search
EmbeddingsStubEmbeddingsService (fixed response)ONNX model inference
Workflow pipelineChatRoute::Workflow stubDagster oneshot job launch
NATS event relayasync-nats in dependencies but not wiredevt.chat.{trace_id} publish
AuthenticationApiKeyAuthorizer defined but not connectedhandler returns Ok(())
PostgreSQL idempotency logsqlx in dependencies but not wiredidempotency_log / completions_cache

Error Format

Errors follow the OpenAI JSON format.

  pub enum ApiError {
    Unauthorized,      // 401
    Forbidden,         // 403
    BadRequest(String), // 400 -> validation_error
    NotFound(String),  // 404 -> validation_error
    Backend(String),   // 502 -> backend_error
    Internal(String),  // 500 -> proxy_error
}
// -> { "error": { "message": "...", "type": "...", "code": N } }
  

Docker Compose Setup

  services:
  proxy:    # Rust proxy :8080
    environment:
      LLM_BASE_URL: http://host.docker.internal:14434  # vLLM/Ollama
      DATABASE_URL: postgres://postgres:postgres@db:5432/openai
      QDRANT_URL: http://qdrant:6333
    depends_on: [db, qdrant]
  qdrant:   # Vector DB :6333
  db:       # PostgreSQL 18 :5432
  

The multi-stage Dockerfile builds with Rust 1.79 and runs on Debian bookworm-slim, with dependency caching.


NATS + Dagster Design Spec (Not Implemented)

This specification was not implemented in Rust. It became the basis for the Go implementation.

Runtime Flow

  1. Client -> /v1/chat/completions(stream=true)
2. Rust allocates trace_id, starts SSE
3. systemd / Quadlet launches dagster-<job>@{trace_id} as oneshot
4. Dagster ops run LLMs and tools in parallel, publish to evt.chat.{trace_id}
5. Rust subscribes -> converts to OpenAI chunks -> streams via SSE
6. On completion, UPSERT to PostgreSQL / Qdrant
  

Event Schema

  {"type":"role","role":"assistant"}
{"type":"token","text":"...","task":"A"}
{"type":"tool_call","name":"search","arguments":"{...}"}
{"type":"usage","usage":{...}}
{"type":"finished","reason":"stop","winner":"llama3.1-8b"}
  

Idempotency Design

  • trace_id: session identifier
  • req_id = sha256(model + messages + params): request identifier
  • All side effects through UPSERTs or unique constraints
  • finished emitted exactly once after artifacts are committed

Staged Migration to JetStream

The plan was to start with NATS Core and add JetStream for durable intake, retries, and replay. The Go version used JetStream from the start.


The Decision to Migrate to Go

The amount of concurrency code led to the Go migration.

Async Context Management Cost

SSE relay, NATS subscriptions, and PostgreSQL/Qdrant writes run within one request. Rust required Pin<Box<dyn Stream>>, tokio::select!, and lifetime management, adding code and debugging work.

Go runs each task in a goroutine and gathers results through channels.

Runtime Overhead

LLM inference I/O dominated this workload. Proxy CPU work was small, so Rust offered little runtime advantage here.

Design Asset Reuse

We reused the Rust type contracts and specifications in Go.

RustGo
trait ChatServiceinterface ChatService
ChatCompletionRequest structChatCompletionRequest struct
ChatRoute enumChatRoute const
ApiError enumApiError type + HTTP status mapping
Layered architectureinternal/transport, internal/domain, internal/infra

Structs, routing, errors, and extension headers carried over.

What Changed After Migration

AreaRust PrototypeGo Production
NATSStubbed (async-nats not wired)JetStream publish (fire-and-forget)
DagsterDesign spec onlydaemon sensor + asset materialization
RAGStubRagServiceKnowledge Service (embed -> pgvector ANN -> rerank)
AuthNot connectedRequestContext middleware with tracking IDs
Telemetrytracing logs onlyVector (Rust) -> Prometheus + Loki + Grafana
Host topologySingle-host Docker Compose3-host (storage / desktop / compute)
RerankerDesign concept onlymulti-bert-inference (Rust + ONNX Runtime) gRPC integration

Caveats

  • The Rust prototype code remains in the openai-api-proxy repository but is not maintained after the Go migration
  • NATS, Dagster, and RAG design specs evolved during the Go implementation (NATS Core to JetStream, oneshot to sensor-driven, Qdrant to pgvector + ColBERT rerank)
  • The Ollama-compatible endpoint was replaced by Anthropic Messages API in the Go version

Verification

  • OpenAI-compatible endpoint confirmed working for both streaming and non-streaming responses
  • Ollama-compatible endpoint DTO conversion verified
  • Multi-backend routing by model name prefix confirmed
  • Docker Compose startup and connectivity for proxy + PostgreSQL + Qdrant verified

Next Steps

The Rust prototype settled the API and type contracts before the Go rewrite.

The Go version runs NATS JetStream, Dagster pull subscriptions and asset materialization, and pgvector ANN + ColBERT rerank RAG. This 3-host AI platform connects local models to application development.

The full design record is in Designing an AI Orchestration Platform with Go, NATS, and Dagster.