As enterprises rapidly integrate artificial intelligence into their daily operations, one of the biggest challenges has become managing the soaring costs associated with running large language models.
Every AI prompt consumes computational resources, often measured in tokens, and as businesses scale AI adoption across thousands of employees, token usage can translate into significant operational expenses.
Against this backdrop, global professional services firm EY has revealed that its internally developed invisible AI router has reduced token consumption by as much as 60%, marking a significant breakthrough in enterprise AI optimization.
Unlike traditional AI systems that send every request to a single large language model, EY’s AI router works behind the scenes, intelligently directing each query to the most suitable model based on the complexity of the task.
Register for Tekedia Mini-MBA edition 20 (June 8 – Sept 5, 2026).
Register for Tekedia AI in Business Masterclass.
Join Tekedia Capital Syndicate and co-invest in great global startups.
Simple requests, such as summarizing documents or answering routine questions, are handled by smaller, less expensive models, while more demanding tasks requiring advanced reasoning are routed to more powerful frontier models. The entire process happens seamlessly, making the routing mechanism effectively invisible to end users.
This intelligent orchestration addresses one of the biggest inefficiencies in enterprise AI deployment. Many organizations rely on premium AI models for every task, regardless of whether such computing power is necessary.
While this guarantees high-quality responses, it also results in excessive token usage and unnecessarily high infrastructure costs. By matching the right model to the right workload, EY has demonstrated that substantial savings can be achieved without compromising user experience.
Token efficiency has become increasingly important as businesses expand AI adoption across departments including finance, legal, consulting, customer support, and software development.
Millions of prompts generated every day can quickly drive cloud computing bills into the millions of dollars annually.
Reducing token consumption by up to 60% represents not only lower operational costs but also improved scalability, enabling organizations to deploy AI more broadly without facing exponential increases in spending.
Beyond financial benefits, the routing system also improves overall performance. Smaller models often generate responses faster than larger ones, reducing latency for routine tasks.
Employees receive quicker answers while organizations reserve premium computing resources for tasks that genuinely require sophisticated reasoning. This balanced allocation enhances productivity and maximizes the return on AI investments.
EY’s approach reflects a broader trend within the AI industry toward multi-model ecosystems. Rather than relying exclusively on a single provider, enterprises are increasingly combining models from different vendors and selecting the best option dynamically.
AI orchestration platforms are becoming essential infrastructure, allowing organizations to balance cost, speed, accuracy, and security according to business requirements. The development underscores a growing shift in enterprise AI strategy.
Competitive advantage is no longer determined solely by access to the most advanced language models but by how intelligently companies manage and optimize those models.
Routing technologies, prompt optimization, caching mechanisms, and workflow automation are emerging as critical tools for improving AI efficiency while controlling expenses.
As AI continues to transform industries, organizations will increasingly prioritize solutions that maximize value rather than simply increasing computing power. EY’s invisible AI router demonstrates that significant efficiency gains can be achieved through smarter system design instead of larger models alone.
The reported reduction in token consumption illustrates how innovation in AI infrastructure can deliver meaningful business outcomes. By optimizing model selection behind the scenes.
EY has shown that enterprises can simultaneously reduce costs, improve performance, and scale AI adoption more sustainably. As businesses continue investing heavily in generative AI.
Intelligent routing technologies are likely to become a standard feature of next-generation enterprise AI architectures, shaping how organizations deploy and manage artificial intelligence in the years ahead.



