Coding teams now have more models and harnesses to choose from than ever. Smart Routing in Unity AI Gateway helps remove that choice overload by matching each coding task to a model, and, with Omnigent, a harness that fits its complexity and cost.
Key highlights
- Route by task, not by default: Smart Routing classifies a task using signals from the initial prompt, then selects a model suited to the work. Simpler tasks can use lower-cost models, while complex tasks can be escalated to more capable options.
- Preserve cache efficiency: The router uses task-aware routing at the start of a session rather than switching models on every request. Keeping a session on the selected model helps protect cache-hit rates and control end-to-end task cost.
- Reduce cost without blunt caps: Internal results showed about 35% savings, while public coding benchmarks showed 56% cost savings. The goal is not to minimize spend at any cost, but to improve productive output per dollar while preserving quality.
- Route across models and harnesses: Smart Routing works with Claude Code and Codex through Unity AI Gateway. With Omnigent, teams can select both the model and coding harness, including for sub-agents.
- Measure productivity, not just price: Coding-session traces provide the feedback loop for evaluating model mix, routed sessions, dollar savings, and developer experience. Routing should be refined using real workloads rather than cost alone.
๐ Read the full post here