Your Gateway Supports the Latest Model. So Does Everyone Else's

Your Gateway Supports the Latest Model

Anthropic released Claude Opus 5.5 on September 22, the first model in its 5.5 family, and Claude Sonnet 5.5 followed six days later. OpenAI spent September expanding GPT-6 with Astra, Sol, and Luna. AI gateway vendors announced support for all of them within days, and the announcements were nearly indistinguishable.

Fast support for a new model is useful, but when every gateway reaches the same models on roughly the same timeline, model support stops being a meaningful differentiator.

Gateways truly start to differ after the model is chosen, in what they send to the model, how much they send, and how tightly they control each request. If you're evaluating AI gateways right now, understanding how each option handles access, tools, data retrieval, and context will tell you far more than which models it supports.

Model breadth followed the same path as connectivity

A model catalog can't differentiate one gateway from another. Provider APIs converge, integrations are cheap to replicate, and every serious gateway reaches the same frontier models within days of a release.

Connectivity went through this first. Supporting the newest data source stopped being a reason to choose a connectivity product years ago. It became the entry requirement, and model access has now followed.

Broad model support still matters. It gives you failover when a provider degrades and flexibility when prices change. But it's a floor, not a ranking. If every option clears the floor, the decision gets made on what the gateway does with a request after the model is chosen. (For a full evaluation framework, see our AI gateway series.)

Ask what an agent is allowed to reach

Start with a more consequential question: Does the gateway scope access to the task, and is that scope enforced where the data actually lives?

An agent should receive tools that expose only the data and operations required for its task. Enforcement at the source means each request carries a delegated identity with the user's real entitlements into the system of record. Writes are then validated against that system's own business rules before they commit.

Without that architecture, teams often fall back to shared service accounts. The result is broader access than security teams want and weaker attribution for actions taken on a user’s behalf.

This is also a cost issue. Broad access encourages broad retrieval, and broad retrieval creates tokens the organization has to pay for. Governance and token efficiency are not competing architectural priorities here. They are consequences of the same design decision.

Ask how much data comes back

Tool definitions used to be the obvious place to look for wasted tokens, and model providers now handle much of that with tool search, which loads only the tools a request needs. What those tools return is a different problem. Does the gateway return the answer, or all the records the answer is hiding in?

A tool that understands the schema, objects, and relationships of the source it reaches can make a precise call, process data at the source, and return the answer set. A connection without that understanding has to retrieve more data and rely on the model to determine what matters.

That distinction affects accuracy as well as cost. When CData Labs tested MCP providers against real enterprise queries, the lower-scoring providers reached the right system but got the meaning wrong. They misread relative dates like "this quarter," misapplied multiple conditions, and read "Highest" as "High."

These mistakes don't return an error, so nothing flags them. The wrong rows go back to the model as input tokens, and they stay in the conversation, which means you pay for them again on every subsequent turn the agent takes.

Unlike semantic caching, prompt compression, or routing to a cheaper model, precise retrieval can reduce token usage without sacrificing answer quality.

Ask who keeps what the gateway learns

The final question plays out over years rather than individual requests: When a definition or a correction is approved, does it stay with your company and travel across models, or does it live inside one provider?

Business context should sit outside any single model or source system. A definition approved once—how the business calculates an active customer, what “pipeline” means to sales, or which segment definition finance uses—should apply to later requests regardless of which model answers them or which source provides the data.

That accumulated meaning is harder to replace than either the underlying model or the systems where the data lives.

The CData Connect AI Gateway Context Engine is built for this. It combines what CData’s connectors know about each system, what your teams have documented, and the institutional knowledge from subject matter experts. The context improves as corrections are approved and stays independent of the model.

Models will change. The meaning a company has built around its data should not have to move with them.

Frequently asked questions

Does model support matter when choosing an AI gateway?

It matters as a floor, not as a differentiator. Every serious gateway reaches the major frontier models within days of a release, so model coverage tells you a product is viable, not that it's the right choice. Confirm breadth for provider failover and the ability to move when prices change, then set it aside.

What should you evaluate in an AI gateway instead of model coverage?

Three things: how narrowly access can be scoped and where that scope is enforced, whether the gateway returns the answer or the full record set, and whether the context it accumulates stays portable across models. The first two carry a direct token cost, which makes them budget questions as well as architecture questions. The third decides whether what your teams approve survives a model change. For a full evaluation framework, see the CData AI gateway series.

Why does an AI gateway affect token cost at all?

Because the gateway decides what reaches the model. Tool definitions loaded upfront and intermediate results passed through the context window both consume tokens on every turn, and a connection that retrieves broadly makes the model sift data it didn't need. Reducing what enters the context is a gateway design decision, not a model selection decision.

Evaluate your gateway on what it sends to the model

With CData Connect AI, you can see what a governed request sends to the model and what it costs. Start a free trial today or book a call to walk through a specific workflow.

Explore CData Connect AI today

See how Connect AI excels at streamlining AI and business processes for real-time insights and action.

Get the trial