Embedding and rerank APIs with Rust and ONNX
A Rust and ONNX retrieval implementation for LLM integration. It separates embeddings and ColBERT reranking and records the API design, initial cache proposal, and switch to direct gRPC vector transfer.
Separate embedding and reranking APIs
I designed embed and rerank APIs using Rust, Axum, ort, and tokenizers. These retrieval components fit into an OpenAI-compatible local LLM stack and RAG integration.
The initial design precomputed document vectors and shared a key -> bin cache over the LAN. ONNX/CPU execution matched the wider proxy and CPU/GPU split.
CPU-side processing
The CPU handles prompt shaping, embeddings, and task normalization; the GPU handles final generation. Embedding and reranking are shared components separate from generation.
ColBERT reranking uses token-level vectors. Sending them repeatedly adds bandwidth and latency, so the initial design transported keys and small metadata while keeping document vectors in cache.
Design Approach
The two contracts are embed(256d) and rerank(ColBERT 64d token, MaxSim), both using Rust and ONNX/CPU.
Precompute document vectors when documents change and generate query vectors per request. The initial key-based cache lets workers share document vectors.
Specification Details
Goals
The embed API returns one vector for nearest-neighbor retrieval. Rerank uses ColBERT late interaction to reorder shortlisted candidates. Each can be tuned separately.
Embed API
POST /embed accepts texts: string[] and returns 256-dimensional vectors. Use tokenizers, ONNX inference through ort, model-specific pooling, truncation to 256 dimensions, and L2 normalization.
A fixed 256-dimensional output simplifies indexing even when the model emits around 1024 dimensions. Pooling and normalization must follow the model guidance, with numerical validation against a Python reference.
Rerank API
POST /rerank takes query: string and candidates: Candidate[], returning scores: float32[] and order: int[]. Candidates use doc_key or key-derivation metadata such as doc_id + seq + chunk_id.
Generate query token vectors (Tq x 64) and fetch document vectors (Td x 64) from cache. MaxSim sums the best document-token match for each query token, preserving ColBERT interaction without sending candidate text each time.
Models and File Layout
The embedding implementation uses lightonai/modernbert-embed-large/onnx/*, with INT8 onnx/model_int8.onnx at roughly 17MB. Quantization helped retrieval precision in this setup; latency was around 10ms. The small model supports parallel execution.
Use tokenizer.json from the same repository and match special tokens and normalization to the reference implementation.
The initial rerank model was mixedbread-ai/mxbai-edge-colbert-v0-32m. If ONNX files are absent, ort cannot run it directly; the alternative is lightonai/mxbai-edge-colbert-v0-32m-onnx.
Cache Design (LAN-based, key/bin)
The cache proposal uses PostgreSQL doc_id + content_seq plus variables affecting the vector payload:
doc_id|content_seq|chunk_id|model_id|tokenizer_hash|max_len|chunk_ver|dtype|layout
Hash the material with blake3, optionally shorten it, and add a document prefix. Advance content_seq only when the document changes. Change chunk_ver to invalidate vectors when chunking changes.
Start with fp32 (per-row float32 vectors) for validation, then consider bitpack (sign encoding + u64 array).
The fixed header contains dtype, dim (=64), n_tokens, scales_present, layout (token-major), and version. MaxSim and decoding must agree on token-major layout.
gRPC Design (Binary Vec Direct Transmission)
During implementation, I removed DB dependency. Instead of sending only doc_keys, the service receives pre-embedded 256d vectors directly.
MaxSimSearch receives map<string, VecF32>: candidate keys paired with 256d vectors. 256 floats x 4 bytes = 1024 bytes/vec, about 1KB per candidate. The server filters with threshold, top_k, and top_p, then returns parallel arrays of keys and scores.
message MaxSimRequest {
map<string, VecF32> candidates = 1; // key -> 256d vec (100-200 typical)
VecF32 query_vec = 2; // query vector (256d)
float threshold = 3; // min score to include
int32 top_k = 4; // max results
float top_p = 5; // cumulative-score cutoff
}
Direct vector transmission suited these conditions:
- Communication is confined to the local LAN
- Each candidate is roughly 1KB; 100-200 candidates stay within 100-200KB
- Removing DB dependency makes the embed/rerank service fully stateless
- Retrieval quality was verified to be sufficient at this precision level
This removed the cache layer and PostgreSQL content_seq management, leaving a stateless service suited to parallel execution.
HTTP /v1/rerank accepts query text and candidate keys and runs ONNX internally. gRPC lets upstream services such as agent-gateway submit existing vectors for MaxSim scoring.
Rust Implementation Requirements
Axum provides REST, with optional internal gRPC. tokenizers loads tokenizer.json and matches reference special tokens and normalization. Begin ort with intra_threads = num_cpus, optimization_level = Level3, and the CPU execution provider.
MaxSim uses inner product by default. A trait boundary allows later testing of bitpack + Hamming approximation without changing request handling or decoding.
Implementation Order
The initial order was /embed with pool -> truncate256 -> normalize, ONNX availability and /rerank, PostgreSQL content_seq, document precomputation and caching, then minimal key-based gRPC.
That order fixes API and array contracts before transport. The later implementation change removed the proposed DB and cache stages.
Caveats and Unknowns
The initial checks were rerank ONNX availability, pooling and normalization fidelity, and invalidation. Changing max_len or stride requires changing chunk_ver in the cache key.
Results
The design separates embedding and reranking so they can run as independent services or within the larger OpenAI-compatible proxy.
The final gRPC implementation sends vectors directly, replacing the initial key-only cache and database-versioning design.
Next Steps
The INT8 embedding model is roughly 17MB with around 10ms latency. Reranking uses lightonai/mxbai-edge-colbert-v0-32m-onnx, which Rust can run directly through ort.
The remaining choice is whether /embed and /rerank are public APIs or internal proxy components. This affects authentication, monitoring, and workflow integration.
