What is model routing?
How to introduce a centralized control plane for routing and monitoring prompt traffic
In this guideHideShow
Model Routing Basics
As the sun sets on the era of tokenmaxxing, companies are taking a fresh look at their AI spending. AI cost optimization is becoming an important topic for business leaders, and model routing is emerging as a lever for managing AI spend without sacrificing productivity and performance.
Model routing is a new AI cost optimization technique that automatically matches prompts to the proper model to optimize cost and performance. Centralized model routing means end users don't have to think about model selection. An appropriate model is automatically assigned to handle each request, factoring in the pre-established expectations around performance and savings. Your routing strategy is personalized to your business and can vary in sophistication based on your needs and goals, ultimately becoming a critical part of your AI infrastructure.
What is a model router?
Systems that handle model routing are unsurprisingly known as model routers. In most modern setups, the model router controls the optimization logic that traffics your requests. They can sit within an AI Gateway, behind the AI Gateway or as a standalone product. The model router processes requests in real-time as they are fed from your system. Routing is designed to be seamless to the requestor, silently distributing requests across a pool of models that you define. The model router is the traffic cop. It decides how each request is processed and coordinates the exchange between the proper provider based on what it determines as the optimal option (the "route").
Model Routing Policies
Model routing strategies are also called the Routing Policy. Strategies are rapidly evolving, ranging from basic rules based routing logic to complex ML-driven routing algorithms. The approach you choose is heavily dependent on your goals. Here are a few things to consider:
- Provider Routing: Services like OpenRouter offer a large collection of hosting providers that enable access to hosted models. One routing approach can focus on distribution across providers based on latency, cost, uptime or model availability.
- Cost Optimization: Not all prompts require the most powerful model available. Routers can use different techniques to select the right model for a given task based on previous performance and cost.
- Geography: If you have a global or distributed business, model routers can move traffic to the proper geographic zone to keep latency low and data storage localized.
These are all useful priorities, but cost optimization is the most common objective. The proliferation of model options combined with runaway AI spend has been the real catalyst for model routers. There are many viable approaches for cost optimization that we lay out below in our section labeled What are some routing strategies.
When does it make sense to introduce dynamic routing?
It's never too early to start thinking about optimization, but your mileage will vary based on your workloads and volume. Here are some rules of thumb to follow. This assumes your goal is to reduce cost, but the advice generalizes to all objectives.
- Meaningful scale: You should assess if you have enough volume to make this exercise worth it. There's no magic threshold, but imagine targeting 20-25% cost improvement as a starting point. Is that meaningful at your current stage?
- Varying prompt complexity: Savings is achieved by routing simple problems to cheaper models. If you are dealing with approximately the same question every time, you're better off finding an optimal model and routing all traffic to that option. Routing becomes useful when prompts have different levels of complexity that you can identify and address in a more cost-effective way
- Aggregation of AI Traffic: Model routers must sit between your end users and your model providers. If your teams are leveraging a system that does not support custom AI endpoints, you won't be able to mediate the inputs and outputs appropriately.
The good news here is that model routing is becoming both a product category and a feature. Even if you don't have meaningful scale yet, it still may make sense to explore an AI gateway. Many AI gateways ship with a model router feature, and this allows you to start with capturing telemetry to evaluate the routing opportunity as you grow.
What's the difference between a model router and an AI gateway? Do I need both?
Model routers and AI gateways are two different categories within the AI infrastructure stack. You can choose to build a custom solution, but most opt for an off-the-shelf provider. There are many viable options for both routing and gateways.
There is a lot of overlap between features, but it's important to think about routers and gateways as distinctly different products. Model routers are strictly focused on optimizing model usage against a stated business objective. Routers make real-time decisions about what models are used for a given request, and can be trained over time to improve. Like most AI products, there's a fair amount of hill-climbing expected in tuning model selection.
AI gateways can be thought of as a more general offering, but they often contain model routing capabilities. Introducing an AI gateway will give your business much better control and visibility, which will ultimately lead to a more secure AI environment for your business. Gateways will generally include more robust observability features such as advanced prompt logging and metadata capture. Gateways use regular expressions to catch leaking API keys, or apply guardrails against dangerous keywords or prompt injection.
A gateway can be valuable without intelligent model selection. For most businesses, an AI gateway is becoming table stakes and it helps to lay a foundation for model routing down the road.
What are some routing strategies?
This section deserves it's own guide, but below is a list of strategies to consider as a starting point if you are exploring how to introduce model routing into your workflow. All of these strategies will involve inspecting the payload being sent by the user, as that's often your only context.
- Rules-Based Routing: The most simplistic approach to routing. A rudimentary router may send short prompts to a cheaper model and long prompts to a more powerful model
- Cascading Routing: This approach involves starting with a cheap model, obtaining a response, evaluating the quality of that response, and then advancing to the next highest model until a minimum quality score is achieved. Latency will be worse with this approach and depending on your model waterfall, it could also be expensive.
- Classifier Based Routing: First, a small model classifies the complexity request, then a rule is applied based on the classification to route the request to the appropriate model.
- ML-Based Routing: A more sophisticated strategy may first rely on machine learning to generate predictions about a request and then use those predictions to feed a policy-based router. These approaches often attempt to predict the outcome of the session, taking into account the user's end goal, not just the current prompt being requested. The ML model may attempt to predict if the session will require tool calls, or how many turns it will take with the user to reach a conclusion.
What is the KV cache and how does it impact model routing?
As model routing evolves, expect to see more approaches that factor in a user's ultimate intent when predicting how to route a request. Session-based optimizations are becoming increasingly common. Early routers evaluated each prompt individually, which worked well on benchmarks but ultimately had muted impact.
Session-based strategies also tend to factor in the benefits gained from caching. You may hear the phrase KV cache, stands for Key-Value cache that exists at the model level. When multiple prompts are sent to the same model with a similar context, the model is able to reference a cached version of previous inference requests to speed up response time and lower computing costs. This means it may be optimal to send a simple request to a more sophisticated model, because it already has a substantial portion of the request cached.
There's a lot of research going into this right now, including how to optimally time changes to route behavior. For example, some routers will re-evaluate model routes whenever compaction is triggered. Compaction is used by models to reduce ("compact") the context on long-running sessions. In this case, the cache benefit is removed because the context is changing, which means the router has an opportunity to update a model preference without incurring an additional cache penalty.
How do I get started with model routing?
If you're ready to start exploring model routers, check out the Catalog of providers through the link below. You can access reviews, some benchmark analysis and compare different options based on what's right for your business.
Continue to the Model Router Catalog→