Work

AI infrastructureOpen-source project2026

kvfleet: cache-aware routing for LLM fleets

A routing control plane that keeps KV-cache locality, policy, and health in one explainable decision.

  1. Normalize

    • Fingerprint

      Conversation and prefix signals

  2. Gate

    • Policy and health

      Hard constraints first

  3. Rank

    • Multi-objective score

      Locality, load, cost

  4. Execute

    • Route with fallback

      Decision explanation attached

From request to an explained route.

Context

Teams running self-hosted or hybrid LLM fleets usually route by model name alone. That throws away the cheapest performance win they have: sending a conversation back to the replica that already holds its KV cache.

The problem

A good route has to respect tenant and compliance policy, endpoint health, cache locality, and cost at the same time, and operators need to see why a route changed when something goes wrong.

My role

Author and maintainer. Designed the routing pipeline, gateway mode, and the decision-explanation model.

Approach

  • Requests are normalized into stable conversation and prefix fingerprints, kept as separate signals because they fail differently.
  • Hard policy and health gates run before any ranking, so an attractive score can never route around an eligibility rule.
  • Eligible endpoints are ranked across locality, load, and cost, and every decision carries a readable explanation.
  • An OpenAI-compatible gateway mode lets existing clients adopt it without code changes.

Outcome

  • Published as a Python package and an OpenAI-compatible gateway.
  • Route decisions are inspectable instead of hidden behind a single score.
  • A long-form architecture article documents the design and its current scaling limits honestly.
Next storyCheckpoint safety for change-data-capture in Apache SeaTunnel