CData Labs
Discover how AI gateway design affects accuracy, token efficiency, and model behavior through research, benchmarks, and analysis from CData Labs.
Model performance on enterprise data
Three tasks against live CRM, warehouse, and ITSM data. R1 asks which customers are at risk, R2 which of those also show declining usage, and A1 writes a review record for every eligible account and nobody else. Ordered by R2 cost per correct answer: the cheapest model and the most expensive return the same answer.
| Model | Tier | R1 Accuracy | R1 $/correct | R2 Accuracy | R2 $/correct | A1 Accuracy | A1 $/run |
|---|---|---|---|---|---|---|---|
| Mistral Small | Economy | 91% | $0.0009 | 100% | $0.0009 | 100% | $0.0014 |
| GPT-5.6 Luna | Economy | 89% | $0.0043 | 99% | $0.0023 | 100% | $0.0047 |
| DeepSeek V4 Flash | Economy | 91% | $0.0021 | 100% | $0.0032 | 100% | $0.0024 |
| Llama 3.3 70B | Economy | 0% | — | 86% | $0.0039 | 0% | $0.0013 |
| Gemini 3.1 Flash-Lite | Economy | 91% | $0.0030 | 100% | $0.0052 | 100% | $0.0052 |
| Grok 4.3 | Mid | 91% | $0.0040 | 100% | $0.0061 | 100% | $0.0067 |
| Mistral Large | Mid | 91% | $0.0068 | 31% | $0.0096 | 100% | $0.0071 |
| GPT-5.4 mini | Economy | 96% | $0.0114 | 98% | $0.0104 | 100% | $0.0098 |
| Grok 4.6 | Frontier | 91% | $0.0235 | 100% | $0.0144 | 100% | $0.0231 |
| Gemini 3.7 Flash | Mid | 100% | $0.0412 | 100% | $0.0156 | 100% | $0.0169 |
| Haiku 4.5 | Economy | 91% | $0.0313 | 100% | $0.0170 | 100% | $0.0256 |
| Qwen 3.7 Max | Mid | 98% | $0.0314 | 100% | $0.0210 | 100% | $0.0254 |
| DeepSeek V4 Pro | Mid | 100% | $0.0519 | 100% | $0.0256 | 100% | $0.0214 |
| GPT-5.6 | Frontier | 0% | — | 100% | $0.0417 | 100% | $0.0879 |
| Gemini 3.5 Flash | Mid | 100% | $0.0645 | 100% | $0.0440 | 100% | $0.0394 |
| GPT-5.5 | Frontier | 89% | $0.0961 | 100% | $0.0568 | 100% | $0.0647 |
| Sonnet 5 | Frontier | 91% | $0.0461 | 100% | $0.0604 | 100% | $0.0531 |
| Sonnet 4.6 | Mid | 82% | $0.1987 | 100% | $0.0686 | 100% | $0.1469 |
| Opus 5 | Frontier | 98% | $0.3274 | 100% | $0.0753 | 100% | $0.3204 |
| Opus 4.8 | Frontier | 95% | $0.4372 | 100% | $0.1450 | 100% | $0.1584 |
| Fable 5 | Frontier | 91% | $0.2436 | 100% | $0.1571 | 100% | $0.2914 |
| Qwen 3.5 9B | Economy | 0% | — | 0% | — | 100% | $0.0043 |
Accuracy is set F1: how closely the returned records match the verified answer, where 100% is an exact match. Cost is per correct answer on the read tasks and per run on the write, where a write either lands correctly or does not. Tier reflects where each developer positions the model in its own lineup, not its price.
Model cost at equal accuracy
Seventeen of 22 models returned the identical correct answer on live CRM, warehouse, and ITSM data. Only the price changed. Based on internal testing by CData Software (Q3 2026). No independent third-party verification. Actual cost gaps varied among models, testing conducted using sandbox accounts containing known data sets that mirror production account structures. Results may not be representative of performance in live production environments, and results may vary. Organizations should conduct their own independent testing before making purchasing or implementation decisions.
Tokens per answered question
The same multi-source question, asked through CData Connect AI instead of raw table access, dropped from 183,541 tokens and 22 tool calls to 4,427 tokens and one. Based on internal testing by CData Software (Q2 2026). No independent third-party verification. Actual token gaps varied among configurations, testing conducted using sandbox accounts containing known data sets that mirror production account structures. Results may not be representative of performance in live production environments, and results may vary. Organizations should conduct their own independent testing before making purchasing or implementation decisions.
Accuracy by server architecture
Five ways of exposing enterprise data to an AI agent, put to 378 prompts across CRM, project management, warehouse, and ERP systems. Small per-step gaps compound fast: at 75% per step, fewer than a quarter of five-step workflows finish correctly. Based on internal testing by CData Software (Q4 2025). No independent third-party verification. Actual accuracy gaps varied among platforms and MCP approaches, testing conducted using sandbox accounts containing known data sets that mirror production account structures. Results may not be representative of performance in live production environments, and results may vary. Organizations should conduct their own independent testing before making purchasing or implementation decisions. 75% range of average accuracy across platforms, results differ by MCP approach.
Open Methodology
Our test harnesses, benchmark runners, and raw results are public. Fork them, challenge our numbers, run them against your own setup.