Comment by leochong
I built Reflex, a GGUF-native Rust + CUDA inference engine optimized for cold-start latency and energy — process launch to first token — instead of sustained server throughput.The core bet: every CUDA kernel is compiled ahead-of-time by nvcc at build time and shipped inside the binary, never compiled at runtime via NVRTC. A naive runtime-JIT design pays a real multi-second tax on first kernel use — we measured vLLM taking ~235s to become ready (CUDA graph capture) before serving a single request. llama.cpp avoids this too, since its kernels are also build-time compiled — the difference is what we then do with a guaranteed-zero-JIT cold start.
Why cold start and not throughput: llama.cpp, vLLM, and Candle all already have multi-year head starts on the steady-state throughput race, and it's a kernel-optimization war Rust-as-a-language doesn't help you win. Cold start — serverless/FaaS, single-shot CLI calls, batch/cron jobs, edge devices that wake on demand — is a comparatively underserved axis, and it rewards different design decisions (e.g. never JIT-compiling anything, ever).
Benchmarks (all cold-start, same RTX A6000, n=3, external wall-clock via /usr/bin/time -v — not our own internal timer):
- vs llama.cpp: ~1.3–1.4x faster (4.71–5.05s vs 6.45–6.56s). Both engines are AOT-compiled, so this isn't the JIT-tax story above — it's just the fast-IO work (weights uploaded once, on-GPU dequant, device-resident activations through a layer). - vs vLLM: ~24–52x faster, but vLLM isn't built for this workload at all — the installed version also had no GGUF support, so this ran against an HF safetensors checkpoint instead (disclosed in the repo, not smoothed over). - vs Ollama: competitive when it doesn't stall (~6–7s), but its bundled llama-server intermittently hits an internal GPU-discovery-watchdog timeout (~55–62s). This mostly measures packaging/daemon overhead, not the AOT-vs-JIT bet. - vs a managed inference API (TypeSafe Jev), cold-start-to-decision: we lose, ~10–60x slower — different deployment model, an always-warm managed API vs. a genuine cold local process launch. Reported as a loss because it is one.
Supports dense Qwen3, Qwen3-MoE, the Qwen3.5 hybrid Gated DeltaNet mixer, and DeepSeek-V2/V3 (Multi-head Latent Attention).
Deliberate non-goals, not just current scope: batch_size is always 1, no request queue, no continuous batching, no in-core HTTP/gRPC server, ever. Multi-tenancy and serving belong in a host orchestrator, not in this engine — the framing is "let vLLM win the warm-throughput race; this wins by being the fastest way to turn cold compute into one output token, then getting out of the way."
MIT licensed, Rust + CUDA, C-FFI and Python bindings included.
https://github.com/lateos-ai/reflex
Feedback welcome, especially on the cold-start energy angle — that's the part of this I think is genuinely underexplored (existing energy benchmarks measure warm/steady-state joules-per-token, not full-lifecycle cold-start cost) and I'd like to hear if others have run into the same gap.