Moving an OpenAI-compatible proxy from Rust to Go
Why a Rust OpenAI/Ollama proxy moved to Go. Covers SSE, routing, RAG, and implementation limits for LLM integration and AI application development.
Working features and the Go migration
I built a Rust (axum) proxy to expose local LLMs through an OpenAI-compatible API. About 1,600 lines handled OpenAI/Ollama clients, SSE streaming, and model-name-based routing. HTTP, Domain, and Infra were separate layers.
NATS relay, Dagster oneshot jobs, RAG, and Qdrant semantic caching remained stubs. Coordinating SSE, subscriptions, and database writes made the Rust implementation costly, so I moved to Go (Gin).
The Rust types and specifications became the basis for the Go implementation. Goroutines and channels made the same control flow easier to write.
This article is the prequel to the full design record of the Go-based AI orchestration platform.
Prerequisites
- Language: Rust 2024 edition
- Framework: axum 0.8, tokio 1.48
- Serialization: serde / serde_json
- HTTP client: reqwest 0.12
- Messaging (design only): async-nats 0.45 (stubbed)
- Data stores (design only): sqlx 0.8 (PostgreSQL), qdrant-client 1.16 (stubbed)
- Logging: tracing / tracing-subscriber
- Containers: Docker (proxy + PostgreSQL 18 + Qdrant)
- Bind address:
0.0.0.0:8080
Design Decisions
Layered Architecture
We separated HTTP, Domain, and Infrastructure according to Clean Architecture dependency rules.
Client (CLI / IDE / API consumer)
|
v
HTTP Layer [bk/http/]
|-- routes.rs endpoint registration
|-- handlers_chat.rs OpenAI/Ollama chat handlers
|-- handlers_embeddings.rs embeddings handler
|-- error.rs OpenAI-compatible error mapping
|
v
Domain Layer [bk/domain/]
|-- chat.rs ChatCompletionRequest/Response, ChatRoute (Direct/Rag/Workflow)
|-- rag.rs RagParams, EmbeddingsRequest/Response
|-- workflow.rs JobRunId
|
v
Service Layer [bk/services/]
|-- ChatService trait -> HttpChatService (prod) / StubChatService (dev)
|-- EmbeddingsService trait -> StubEmbeddingsService
|-- RagService trait -> StubRagService
|
v
Infrastructure Layer [bk/infra/]
|-- llm.rs HttpClient (backend LLM calls)
|-- qdrant_client.rs QdrantClient (vector search)
|-- auth.rs ApiKeyAuthorizer
Domain has no external dependencies. Backend and storage changes stay in Infra. The Go version kept these boundaries.
Dual OpenAI/Ollama Compatibility
One entry point serves OpenAI and Ollama clients.
| Method | Path | Format |
|---|---|---|
| POST | /v1/chat/completions | OpenAI compatible |
| POST | /v1/embeddings | OpenAI compatible |
| GET | /v1/models | OpenAI compatible |
| POST | /api/chat | Ollama compatible |
| GET | /api/tags | Ollama compatible |
| GET | /health | Health check |
| GET | /ready | Readiness |
The Ollama handler converts DTOs before calling the shared ChatService.
Multi-Backend Routing
The proxy selects a backend automatically based on model name prefix.
// HttpChatService
fn pick_client(&self, model: &str) -> &HttpClient {
// model prefix -> route-specific client, fallback -> default client
}
Configured route clients are matched by model name; unmatched requests use the default. Routing stays in Service rather than HTTP.
ChatRoute Branching
Custom headers on the request explicitly select the processing path.
pub enum ChatRoute {
Direct, // forward to backend LLM
Rag, // vector search + context injection + LLM
Workflow, // delegate to Dagster pipeline
}
x-workspace identifies a RAG workspace, x-pipeline a workflow, and x-correlation-id a trace. These headers express direct inference, RAG, and workflows through /v1/chat/completions.
SSE Streaming
The service layer returns one of three response types.
pub enum ChatServiceResponse {
Once(ChatCompletionResponse), // non-stream: single JSON
Stream(Pin<Box<dyn Stream<...>>>), // chunked: OpenAI-compatible
StreamRaw(Pin<Box<dyn Stream<...>>>), // raw text: backend passthrough
}
Stream mode sends data: {...}\n\n SSE frames and terminates with [DONE]. StreamRaw passes through the backend SSE response without transformation.
Implementation
What Worked
The following prototype functions worked.
/v1/chat/completions— OpenAI-compatible streaming and non-streaming responses/api/chat— Ollama compatibility with internal DTO conversion/v1/models— model listing from backends (both OpenAI and Ollama formats)- Multi-backend routing — automatic client selection by model name prefix
- SSE streaming — parsing SSE lines from reqwest byte streams and relaying them
- Structured logging — tracing spans per handler recording model name, stream flag, and message count
What Remained Stubbed
These functions had traits and specifications but remained stubs.
| Feature | Status | Design Position |
|---|---|---|
| RAG context building | StubRagService (returns fixed string) | Qdrant + PostgreSQL pgvector search |
| Embeddings | StubEmbeddingsService (fixed response) | ONNX model inference |
| Workflow pipeline | ChatRoute::Workflow stub | Dagster oneshot job launch |
| NATS event relay | async-nats in dependencies but not wired | evt.chat.{trace_id} publish |
| Authentication | ApiKeyAuthorizer defined but not connected | handler returns Ok(()) |
| PostgreSQL idempotency log | sqlx in dependencies but not wired | idempotency_log / completions_cache |
Error Format
Errors follow the OpenAI JSON format.
pub enum ApiError {
Unauthorized, // 401
Forbidden, // 403
BadRequest(String), // 400 -> validation_error
NotFound(String), // 404 -> validation_error
Backend(String), // 502 -> backend_error
Internal(String), // 500 -> proxy_error
}
// -> { "error": { "message": "...", "type": "...", "code": N } }
Docker Compose Setup
services:
proxy: # Rust proxy :8080
environment:
LLM_BASE_URL: http://host.docker.internal:14434 # vLLM/Ollama
DATABASE_URL: postgres://postgres:postgres@db:5432/openai
QDRANT_URL: http://qdrant:6333
depends_on: [db, qdrant]
qdrant: # Vector DB :6333
db: # PostgreSQL 18 :5432
The multi-stage Dockerfile builds with Rust 1.79 and runs on Debian bookworm-slim, with dependency caching.
NATS + Dagster Design Spec (Not Implemented)
This specification was not implemented in Rust. It became the basis for the Go implementation.
Runtime Flow
1. Client -> /v1/chat/completions(stream=true)
2. Rust allocates trace_id, starts SSE
3. systemd / Quadlet launches dagster-<job>@{trace_id} as oneshot
4. Dagster ops run LLMs and tools in parallel, publish to evt.chat.{trace_id}
5. Rust subscribes -> converts to OpenAI chunks -> streams via SSE
6. On completion, UPSERT to PostgreSQL / Qdrant
Event Schema
{"type":"role","role":"assistant"}
{"type":"token","text":"...","task":"A"}
{"type":"tool_call","name":"search","arguments":"{...}"}
{"type":"usage","usage":{...}}
{"type":"finished","reason":"stop","winner":"llama3.1-8b"}
Idempotency Design
trace_id: session identifierreq_id = sha256(model + messages + params): request identifier- All side effects through UPSERTs or unique constraints
finishedemitted exactly once after artifacts are committed
Staged Migration to JetStream
The plan was to start with NATS Core and add JetStream for durable intake, retries, and replay. The Go version used JetStream from the start.
The Decision to Migrate to Go
The amount of concurrency code led to the Go migration.
Async Context Management Cost
SSE relay, NATS subscriptions, and PostgreSQL/Qdrant writes run within one request. Rust required Pin<Box<dyn Stream>>, tokio::select!, and lifetime management, adding code and debugging work.
Go runs each task in a goroutine and gathers results through channels.
Runtime Overhead
LLM inference I/O dominated this workload. Proxy CPU work was small, so Rust offered little runtime advantage here.
Design Asset Reuse
We reused the Rust type contracts and specifications in Go.
| Rust | Go |
|---|---|
| trait ChatService | interface ChatService |
| ChatCompletionRequest struct | ChatCompletionRequest struct |
| ChatRoute enum | ChatRoute const |
| ApiError enum | ApiError type + HTTP status mapping |
| Layered architecture | internal/transport, internal/domain, internal/infra |
Structs, routing, errors, and extension headers carried over.
What Changed After Migration
| Area | Rust Prototype | Go Production |
|---|---|---|
| NATS | Stubbed (async-nats not wired) | JetStream publish (fire-and-forget) |
| Dagster | Design spec only | daemon sensor + asset materialization |
| RAG | StubRagService | Knowledge Service (embed -> pgvector ANN -> rerank) |
| Auth | Not connected | RequestContext middleware with tracking IDs |
| Telemetry | tracing logs only | Vector (Rust) -> Prometheus + Loki + Grafana |
| Host topology | Single-host Docker Compose | 3-host (storage / desktop / compute) |
| Reranker | Design concept only | multi-bert-inference (Rust + ONNX Runtime) gRPC integration |
Caveats
- The Rust prototype code remains in the
openai-api-proxyrepository but is not maintained after the Go migration - NATS, Dagster, and RAG design specs evolved during the Go implementation (NATS Core to JetStream, oneshot to sensor-driven, Qdrant to pgvector + ColBERT rerank)
- The Ollama-compatible endpoint was replaced by Anthropic Messages API in the Go version
Verification
- OpenAI-compatible endpoint confirmed working for both streaming and non-streaming responses
- Ollama-compatible endpoint DTO conversion verified
- Multi-backend routing by model name prefix confirmed
- Docker Compose startup and connectivity for proxy + PostgreSQL + Qdrant verified
Next Steps
The Rust prototype settled the API and type contracts before the Go rewrite.
The Go version runs NATS JetStream, Dagster pull subscriptions and asset materialization, and pgvector ANN + ColBERT rerank RAG. This 3-host AI platform connects local models to application development.
The full design record is in Designing an AI Orchestration Platform with Go, NATS, and Dagster.
