How Pareto routing works
New routing strategies emerge that seek a balance around quality, cost, and latency.
In this guideHideShow
What is Pareto Model Routing?
Pareto model routing is an advanced model routing strategy that attempts to balance the trade offs between different models to achieve an optimal outcome. This is most commonly used when balancing the trade off between price and answer quality. The Pareto approach is a push for value. When executed well, the reduction in price is steeper than the reduction in performance.
When you hear discussions about Pareto, you often hear terms such as Pareto frontier or Pareto dominance, two concepts that help evaluate the trade off you are making by swapping models with various performance.
- The Pareto frontier is the boundary where improvement in one of your evaluated criteria comes at a trade off of another metric.
- Pareto dominance refers to models that fall below the Pareto Frontier, meaning there is another model choice that delivers superior results on at least one metric while also delivering equivalent or better results on all other criteria
Let's apply these concepts to a common example around performance and cost. Imagine you feed various models a task and evaluate the performance of the response across each model. At the end you would have the cost and a derived quality score, which you could plot on their respective axes similar to the chart below.

The blue dots on this chart represent the Pareto frontier. If you were to use the model at Point D for example, you would generate the optimal performance, but incur a higher cost than A, B or C. You can see from the chart, that any change in price from Point D also introduces a drop in quality. You are at the frontier, because every change in one axis introduces a trade off with another metric.
Now take a look at the grey dots on this chart. Those dots are all said to be dominated. At each of the grey dots, you can improve both price and performance by choosing a different model. The models that represent the dominated plots would never be an optimal choice in this scenario, unless there was a decision criteria not represented by the chart.
When Is Pareto Model Routing a Good Option?
Pareto Model Routing is most often used in coding use cases, and many models have been tuned specifically to accommodate a sliding scale of complexity. That said, it's important to have a large disparity or distribution of prompt complexity across payloads. If all requests are about the same in terms of complexity, there's limited utility for model routing. A larger distribution gives the model options for trading off cost or speed vs answer quality, which over time will optimize and turn into improvement over blind model selection. If you aren't sure if you have the distribution you need, log a few days worth of prompts and test offline to validate complexity.
Ultimately Pareto is looking for scenarios where a cheaper model gives about the same performance. Savings is achieved by finding the flat part of the curve on the Pareto frontier. For less complex tasks, the most capable model may be providing a negligible improvement in performance at a substantial cost premium. In that scenario, downgrading the model is likely a worthwhile trade off.
Consider the following example:

Is Model F or G ever really worth it in this example? It depends largely on your use case and tolerance for changes in performance. In many circumstances though, Model E or even D provide great value benefits. A significant reduction in cost is realized with no meaningful change in performance.
Growing Support for Pareto Model Routing
As both coding agents and Pareto model routing become more mainstream, support for Pareto out-of-the-box is becoming more common. I'm in the process of evaluating router options, but so far have found that OpenRouter has a clean turnkey option. It can be configured through their admin menu if you want to try this in action.
This is a great opportunity to optimize as new models get released, or as pricing changes. Take this last final example to demonstrate how price changes and model releases shift the frontier.

In this scenario we've removed any loss in performance for a reduction in price from points E to G. Not only does this make the choice even more obvious, but it also removes F and G from the Pareto frontier. There is now a scenario where you can improve a metric (cost) without a trade off, so F and G become grey dots in our theoretical example.
Pareto In Action with Grok
Ongoing improvements in model performance will continue to surface new opportunities for Pareto optimization. Some notable AI investors did a Pareto frontier breakdown on X after the Grok 4.6 release.
This tweet from Gavin Baker plots Grok 4.6 against other models on price and performance. You can see that Grok 4.6 appears to perform roughly on par with Fable, while costing significantly less. It's another example of a "flat" frontier, meaning it's a potentially worthwhile trade off to migrate to Grok 4.6.
This is based on some set of example prompts, so you need to test these performance thresholds on your own workloads. That said, if certain types of prompts produce similar results across both models, those requests can be routed to Grok 4.6. Fable can then be used only where the additional performance is worth paying for. Since performance matters to everyone, this is generally the shape of the savings for model routing - roughly flat performance at a lower cost.