Now: a native AI engine in Rust
Rebuilding the core inference engine in Rust, over a pinned llama.cpp
- Why
- The apps' AI server ran its inference loop through several layers. I wanted one fast, owned engine that every app can switch to without code changes.
- Built
- A Rust server (tokio, axum) over pinned, checksum-verified llama.cpp builds for CPU, Vulkan and CUDA: request scheduling, chat templates, embeddings, resumable model downloads with SHA-256 checks, a supervisor that restarts the engine, and an OpenAI-compatible API. Around it: an admin API and web console, licence-gated HTTPS, a Windows service, metrics and a Linux container image.
- Results
- Same laptop, same session, small models (1.5B and 3B), against my existing AI server: about 1.5× faster generation, first token in 45 ms instead of 163 ms, 2.2–3.1× more throughput with four clients at once, and a cold start answering in 0.8 s instead of 7.4 s. It passes 30 of 30 API conformance checks, and existing apps connect to it with no code change (5 of 5 end-to-end tests).
- Trade-offs
- It uses more memory while a model is loaded, and it has no audio, vision or governance features yet. I measured those costs too, and decided to go ahead for text models.