AI infrastructureOpen-source project2026
kvfleet: cache-aware routing for LLM fleets
A routing control plane that keeps KV-cache locality, policy, and health in one explainable decision.
Normalize
Fingerprint
Conversation and prefix signals
Gate
Policy and health
Hard constraints first
Rank
Multi-objective score
Locality, load, cost
Execute
Route with fallback
Decision explanation attached
Context
Teams running self-hosted or hybrid LLM fleets usually route by model name alone. That throws away the cheapest performance win they have: sending a conversation back to the replica that already holds its KV cache.
The problem
A good route has to respect tenant and compliance policy, endpoint health, cache locality, and cost at the same time, and operators need to see why a route changed when something goes wrong.
My role
Author and maintainer. Designed the routing pipeline, gateway mode, and the decision-explanation model.
Approach
- Requests are normalized into stable conversation and prefix fingerprints, kept as separate signals because they fail differently.
- Hard policy and health gates run before any ranking, so an attractive score can never route around an eligibility rule.
- Eligible endpoints are ranked across locality, load, and cost, and every decision carries a readable explanation.
- An OpenAI-compatible gateway mode lets existing clients adopt it without code changes.
Outcome
- Published as a Python package and an OpenAI-compatible gateway.
- Route decisions are inspectable instead of hidden behind a single score.
- A long-form architecture article documents the design and its current scaling limits honestly.