Skip to content
Tiatra, LLCTiatra, LLC
Tiatra, LLC
Information Technology Solutions for Washington, DC Government Agencies
  • Home
  • About Us
  • Services
    • IT Engineering and Support
    • Software Development
    • Information Assurance and Testing
    • Project and Program Management
  • Clients & Partners
  • Careers
  • News
  • Contact
 
  • Home
  • About Us
  • Services
    • IT Engineering and Support
    • Software Development
    • Information Assurance and Testing
    • Project and Program Management
  • Clients & Partners
  • Careers
  • News
  • Contact

Snowflake adds dynamic model routing to Cortex AI Gateway to cut enterprise AI costs

Snowflake on Tuesday unveiled a dynamic model routing capability for its Cortex AI Gateway, designed to help enterprises reduce AI spending by automatically directing workloads to the most appropriate model based on cost, performance, and latency requirements.

The new capability, which is expected to be in private preview soon, will allow enterprises to define which models they approve for use and the tradeoffs they want the system to prioritize, such as cost, performance, and latency, for an individual application or workload, CEO Sridhar Ramaswamy wrote in a blog post.

Once those policies are defined, Cortex AI Gateway then evaluates each task against those policies and real-world model performance and cost data to determine which model should handle the workload in the most efficient manner, Ramaswamy added.

Further, the CEO pointed out that Cortex AI Gateway also creates a feedback loop by evaluating the quality of a model’s output after it completes a task, which allows the routing system to adjust its decisions as model capabilities, pricing, and performance change, with the aim of continuously optimizing the balance between quality, cost, and latency.

According to Snowflake’s internal benchmarks, the new capability can improve token efficiency compared with using a frontier model for every task.

In one internal test, agents using dynamic routing built a dbt pipeline with up to three times greater token efficiency than a frontier-model-only approach while maintaining the same quality, the company said in a statement. In another test, engineering teams completed the same number of pull requests with 25% greater token efficiency, it added.

Routing could lower AI costs, but adds governance complexity

The new capability will have the largest impact on high-volume, low-complexity workloads where many requests do not require frontier-model reasoning, like classification, extraction, summarization, routine data engineering, and repetitive agent steps, said Stephanie Walter, practice lead of AI stack at HyperFRAME Research. “Routing those requests to smaller models could materially reduce inference costs while preserving expensive models for genuinely difficult tasks,” said Stephanie Walter, practice lead of AI stack at HyperFRAME Research.

Agentic applications could specifically benefit from dynamic model routing, said Advait Patel, senior site reliability engineer (SRE) at Broadcom.

“An agent loop spends most of its steps on plumbing, reading a file, parsing a result, and picking the next call. Very few of those need deep reasoning, but they all hit the same model today. When I pulled telemetry on our own coding agent usage, the spend wasn’t in the hard problems at all. It was the volume of ordinary calls,” Patel said.

However, Walter cautioned that enterprises should not treat token efficiency as the same as cost savings, especially in agentic applications, despite Snowflake’s “promising” internal benchmarks.

“Enterprises must also measure retries, failed tasks, latency, human correction, and the cost of operating the routing layer,” Walter noted.

More so because routing, despite removing the repetitive model-selection work from individual applications, shifts operational complexity into the orchestration and governance layer and doesn’t eliminate it completely, according to Phil Fersht, CEO of HFS Research.

“Enterprises would still need to determine which models are approved, establish routing policies, monitor quality, control costs, and manage security and compliance,” Fersht said, adding that if policies are not defined well, the system can make a poor decision, which at scale, could either produce inconsistent outcomes or unnecessary costs.

That shift of operational complexity into the governance layer, according to Manoj Chandra Jha, principal analyst at Nord-IQ Research, could be challenging for most enterprises: “Short-term complexity can rise, since most teams lack the governance and monitoring maturity routing now requires.”

Routing also adds a new variable for developers to track

The governance burden also has implications for developers, who will have to account for routing decisions as another variable when building and troubleshooting applications.

“Dynamic routing makes visibility essential. If different requests go to different models, developers need to know which model handled a request, why it was selected, and whether the result met expected quality and performance levels,” said Robert Kramer, managing partner at KramerERP.

“When something breaks, they need to determine quickly whether the fault came from the application, the model, or the routing decision. That third failure mode is new, and it is the one teams are least equipped to diagnose today,” Kramer added.

Snowflake, however, is looking to address concerns around changes in routing decisions driven by model pricing changes.

It would integrate Cortex AI Gateway with its AI coding assistant CoCo’s existing role-based access and tagging framework, which will allow enterprise administrators to set default models, attribute AI usage to teams or cost centers, establish per-user quotas, and receive alerts as consumption approaches predefined limits.

These controls could help enterprises maintain visibility into how routing decisions affect AI spending as models, pricing, and workloads change, the company said.

Model routing becomes a new battleground in the AI stack

That enterprise focus on controlling AI spending via model selection and routing hasn’t escaped the attention of other vendors.

Nvidia has been expanding its efforts around model routing, while Cloudflare and OpenRouter have also emerged as players in the space, reflecting growing interest in helping enterprises route workloads across multiple models based on factors such as cost, performance, and capability.

The shift, according to Fersht, is partly a consequence of the growing number of models available to enterprises and the differences between them in cost, performance, latency, and capabilities.

That shifts the strategic value towards the layer that decides which model to use and orchestrates it across enterprise workflows, Fersht noted.

However, Patel cautioned that enterprises should evaluate model routers based on the level of control and transparency they provide.


Read More from This Article: Snowflake adds dynamic model routing to Cortex AI Gateway to cut enterprise AI costs
Source: News

Category: NewsAugust 18, 2026
Tags: art

Post navigation

PreviousPrevious post:The AI leapfrog effect: Why standardizing now could be your costliest mistakeNextNext post:Not every problem needs an AI agent

Related posts

Your identity governance wasn’t built for AI agents
August 21, 2026
Inside TIAA’s massive IT transformation to fuel business growth
August 21, 2026
The decision line
August 21, 2026
Ransomware takes aim at enterprise resilience
August 21, 2026
The more efficient AI makes us, the more human we must become
August 21, 2026
Graph engineering is where AI agents stop working alone
August 20, 2026
Recent Posts
  • Your identity governance wasn’t built for AI agents
  • Inside TIAA’s massive IT transformation to fuel business growth
  • The decision line
  • Ransomware takes aim at enterprise resilience
  • The more efficient AI makes us, the more human we must become
Recent Comments
    Archives
    • August 2026
    • July 2026
    • June 2026
    • May 2026
    • April 2026
    • March 2026
    • February 2026
    • January 2026
    • December 2025
    • November 2025
    • October 2025
    • September 2025
    • August 2025
    • July 2025
    • June 2025
    • May 2025
    • April 2025
    • March 2025
    • February 2025
    • January 2025
    • December 2024
    • November 2024
    • October 2024
    • September 2024
    • August 2024
    • July 2024
    • June 2024
    • May 2024
    • April 2024
    • March 2024
    • February 2024
    • January 2024
    • December 2023
    • November 2023
    • October 2023
    • September 2023
    • August 2023
    • July 2023
    • June 2023
    • May 2023
    • April 2023
    • March 2023
    • February 2023
    • January 2023
    • December 2022
    • November 2022
    • October 2022
    • September 2022
    • August 2022
    • July 2022
    • June 2022
    • May 2022
    • April 2022
    • March 2022
    • February 2022
    • January 2022
    • December 2021
    • November 2021
    • October 2021
    • September 2021
    • August 2021
    • July 2021
    • June 2021
    • May 2021
    • April 2021
    • March 2021
    • February 2021
    • January 2021
    • December 2020
    • November 2020
    • October 2020
    • September 2020
    • August 2020
    • July 2020
    • June 2020
    • May 2020
    • April 2020
    • January 2020
    • December 2019
    • November 2019
    • October 2019
    • September 2019
    • August 2019
    • July 2019
    • June 2019
    • May 2019
    • April 2019
    • March 2019
    • February 2019
    • January 2019
    • December 2018
    • November 2018
    • October 2018
    • September 2018
    • August 2018
    • July 2018
    • June 2018
    • May 2018
    • April 2018
    • March 2018
    • February 2018
    • January 2018
    • December 2017
    • November 2017
    • October 2017
    • September 2017
    • August 2017
    • July 2017
    • June 2017
    • May 2017
    • April 2017
    • March 2017
    • February 2017
    • January 2017
    Categories
    • News
    Meta
    • Log in
    • Entries feed
    • Comments feed
    • WordPress.org
    Tiatra LLC.

    Tiatra, LLC, based in the Washington, DC metropolitan area, proudly serves federal government agencies, organizations that work with the government and other commercial businesses and organizations. Tiatra specializes in a broad range of information technology (IT) development and management services incorporating solid engineering, attention to client needs, and meeting or exceeding any security parameters required. Our small yet innovative company is structured with a full complement of the necessary technical experts, working with hands-on management, to provide a high level of service and competitive pricing for your systems and engineering requirements.

    Find us on:

    FacebookTwitterLinkedin

    Submitclear

    Tiatra, LLC
    Copyright 2016. All rights reserved.