By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Salesforce VP: Enterprise AI Costs Need Architectural Fixes

Enterprises have shifted their focus from rapid AI adoption to controlling costs, moving beyond questions of "how much AI" or "which model is smartest" to "how to generate impactful outcomes while keeping AI costs under control." This change in perspective reflects a realization that the initial maximalist approach to AI deployment, characterized by launching numerous use cases, onboarding many employees, and deploying abundant agents, has led to unsustainable token consumption. Companies like Uber have already tightened internal AI spending after consumption exceeded plans, indicating a growing concern over runaway costs.
While AI providers such as Anthropic, OpenAI, and Google have responded by aggressively cutting input token prices and introducing prompt caching features, and tools like Cursor with Composer 2.5 are now factoring cost into model selection alongside performance, these measures are seen as insufficient. A Salesforce Vice President highlights that these solutions address only a symptom, not the root cause. The core issue lies in the fundamental architecture of most enterprise AI pipelines, which are inefficient and "leak" tokens and associated costs at various execution phases. The analogy of a sieve is used to describe this phenomenon, suggesting that simply making tokens cheaper will not resolve the problem if the underlying system is inherently wasteful.
To achieve cost-effective and impactful AI outcomes, companies must address these inefficiencies through architectural redesign. The current AI pipeline is described as having five predictable failure points where tokens and money are lost. The first identified leak is "Context Bloat," which occurs when excessive information is included in prompts. This not only increases costs but also leads to model drift, diminishing the quality of the AI's output. Addressing context bloat requires more than just sophisticated prompting techniques; it necessitates a foundational architectural approach designed to optimize token usage throughout the AI execution process.
Original source — read the full reporting at the publisher:
Read on FortuneGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.