vLLM Semantic Router: Mixture-of-Models Routing
GitHub Repo
Apache-2.0
October 2, 2026 at 09:19 AM
0 views

vLLM Semantic Router: Mixture-of-Models Routing

@vllm-projectProject Author

What vLLM Semantic Router is

vLLM Semantic Router (often shortened to vLLM SR) is a programmable routing layer for building Mixture-of-Models systems across heterogeneous LLM infrastructure. Its README describes the job plainly: evaluate request signals, user preferences and application policies, then select or compose the right model path for each request. The tagline is Make Your Mixture-of-Models Programmable.

The problem it targets is familiar to anyone running more than one model in production. Different requests need different things: a fast local model, a specialist or frontier model, retrieval, memory, tools, a verifier, or several models working together. Those options carry different trade-offs in capability, latency, cost and trust, and the right choice can change by user, session and available capacity. When every application hard-codes those choices, the project's docs argue, product code becomes coupled to the current model fleet and the same routing logic is repeated across clients. Semantic Router moves that decision into a shared layer in the request path, so applications keep calling one stable OpenAI- or Anthropic-compatible endpoint while the model path behind it evolves.

The intended users are platform and infrastructure teams who operate several models or providers and want routing to be policy-driven, explainable and changeable without application deploys. It lives under the vllm-project organization on GitHub.

How it works

The system overview splits the project into a data plane and a control plane.

Data plane

  • Envoy accepts client traffic, calls the Router through Envoy's External Processing (ExtProc) protocol, and forwards the resulting request upstream.
  • Semantic Router extracts signals, evaluates policy, applies route-specific behaviour, and selects or coordinates model candidates.
  • Backends are OpenAI-compatible model services or provider endpoints. The Router does not load model weights itself.

Control plane

Routing behaviour is defined in canonical YAML. A CLI (vllm-sr) and a dashboard support setup, validation, model discovery and configuration, while Helm charts and an Operator deploy the Router into Kubernetes. Metrics, replay and evaluation expose route outcomes so operators can test and refine policy.

Core objects

  • Entrypoint: maps one or more public model aliases to a recipe.
  • Recipe: a complete routing policy and runtime-state isolation boundary.
  • Signal: a named fact about the request, identity, conversation or content.
  • Projection: a reusable score, partition or band derived from signals.
  • Decision: a policy rule that chooses an eligible route and candidate set.
  • Plugin: route-specific processing such as request controls, memory, retrieval or response handling.
  • Algorithm: the method used to select or coordinate candidate models.
  • Provider model: a physical inference endpoint available to one or more recipes.

A request flows like this: a client sends OpenAI Chat Completions, OpenAI Responses or Anthropic Messages; Envoy hands it to the Router; the requested model name resolves to an entrypoint and recipe; the Router extracts signals and computes projections; decisions pick an eligible candidate set; the route's algorithm selects one model or runs a bounded multi-model strategy; plugins run at their hooks; and Envoy forwards the provider-shaped request and returns a normalized response. The selected model is written to an x-selected-model header. Requests that name a physical model directly pass straight through without signals, decisions or plugins, and if no decision matches inside a recipe, a configured default model is used.

Key features

  • Signal-driven routing: decisions can consider intent, difficulty, context, modality, identity, risk, preference and system state rather than a static model name.
  • Virtual models: entrypoints and recipes turn a shared model pool into purpose-built model aliases that clients call by objective.
  • Multi-model execution: a recipe can pick one model, escalate through a cascade, or coordinate a bounded multi-model workflow. Recent project blog posts cover fusion and collaboration between models.
  • Attached capabilities: retrieval, memory, tool filtering, caching, safety checks and verification can be attached per route.
  • Protocol translation: client and backend wire formats do not have to match; the docs include a compatibility matrix.
  • Deployment flexibility: the same routing model applies in local Docker, Kubernetes and hybrid environments.
  • Evidence and evaluation: routing metadata, feedback, replay and evaluation workflows show what happened and why.

Getting started

The README's install command uses the project's install script on the stable channel:

curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel stable

The installation guide on vllm-sr.ai covers pip, uv and agent-driven installation, plus Docker, Kubernetes and hardware-specific paths. Before installing anything, there is a hosted playground at app.vllm-sr.ai/playground; the README publishes read-only demo credentials for it. Piping a remote script into bash is convenient for a laptop trial, but read the script first or use the package-manager route on shared machines.

From there the docs suggest a path: run the quickstart and send a request through the Router, read the system overview and routing pipeline pages, then define entrypoints and recipes for your own model pool.

Use cases

  • Cost and quality tiering: send simple requests to a small, cheap model and escalate harder ones to a frontier model, without the client choosing.
  • Data boundaries: keep sensitive requests on private or edge infrastructure while routing others to cloud providers, matching the README's goal of keeping data within its boundaries.
  • Specialist routing: send domain-specific prompts to specialist models behind a single public model alias.
  • Safety and verification: attach safety checks or a verifier to specific routes rather than to every call. The project has published work on real-time hallucination detection.
  • Reasoning control: one of the project's papers, When to Reason, deals with deciding when reasoning modes are worth their cost.
  • Heterogeneous hardware: route across GPUs, accelerators, edge devices and cloud providers in one policy.

How it compares

The project's own docs are unusually explicit about where it sits. They describe three routing layers that may all appear in one stack:

  • AI gateway: client ingress, provider translation, credentials, rate limits and traffic policy. Examples named in the docs are Agent Router (formerly Envoy AI Gateway), LiteLLM and agentgateway.
  • Semantic Router: choosing a logical model or model pool from request intent and policy.
  • Inference router: picking a healthy replica inside the selected pool. Examples named are llm-d, vLLM Router and the AIBrix gateway.

So vLLM SR is not a replacement for an AI gateway or for inference schedulers; it is the semantic decision layer between them, and the docs note that Agent Router and agentgateway call it through ExtProc. Teams already using a gateway with simple fallback rules may find that sufficient; vLLM SR becomes interesting when routing depends on request content and policy rather than availability alone. It also does not provision models or manage capacity; that stays with the backend platform.

Things to know before adopting

  • Envoy dependency: the data plane is built around Envoy and ExtProc, so operating it means operating Envoy, either directly or through a supported gateway.
  • Moving quickly: the project was created in August 2025 and has shipped v0.1 Iris (January 2026), v0.2 Athena (March 2026) and v0.3 Themis (June 2026). Configuration concepts have evolved across releases; pin versions and read release notes when upgrading.
  • Research-heavy: the project publishes papers, a white paper and a vision paper alongside code. That is useful context, but separate research ideas from what a given release supports.
  • Policy design is the real work: the Router makes routing programmable, but someone still has to define signals, decisions and recipes and evaluate them against real traffic.
  • Sponsorship: the README credits AMD with GPU resources and ROCm software for router model training, end-to-end testing and the online playground.

Project activity

As of October 2026 the repository has roughly 6,000 stars. It was created on 26 August 2025, is written primarily in Go, and is licensed under Apache-2.0. The source is at github.com/vllm-project/semantic-router and documentation, blog and publications live at vllm-sr.ai. The community meets in the #semantic-router channel on vLLM Slack and holds two monthly community meetings, one timed for APAC and one for the Americas.

Enjoying this project?

Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.

Project
vllm-semantic-router
Created
October 2
Last Updated
October 2, 2026 at 09:19 AM

Find more projects like this

One email a week: new and trending developer tools, fresh comparisons, and what shipped. Unsubscribe in one click.