Abraham Sanieoff has been closely following one of the most fascinating paradoxes in the technology world right now. Artificial intelligence is getting dramatically cheaper to access, yet businesses are somehow spending more on it than ever before. At first glance, that sounds like a contradiction. In practice, it is one of the most important economic patterns defining how companies will compete, operate, and grow over the next several years. If you want to understand where AI spending is actually headed, and why the cheapest models are not automatically producing the lowest bills, this breakdown from Abraham Sanieoff is exactly where to start.
The story begins with a number that deserves to stop you in your tracks. Academic research published in the Journal of Economic Perspectives in 2026 found that the price of AI intelligence has fallen roughly one thousand times over a relatively short period. Open-source models can now cost approximately 90 percent less than comparable closed-source alternatives. The number of commercially available models and inference providers has expanded dramatically. By nearly every measure, the raw cost of accessing capable artificial intelligence is collapsing at a speed that has few historical parallels in any industry.
And yet, the AI bill is not shrinking. For many organizations, it is growing. Understanding why that is happening, and what it means for how businesses should think about AI investment, is the central question Abraham Sanieoff wants to explore here.
The Inference Paradox That Is Reshaping AI Economics
The first generation of generative AI economics was relatively simple. A human asked a question, an AI generated a response, and tokens were billed accordingly. The math was easy to follow. You could look at a price per million tokens, estimate your usage, and project a monthly cost with reasonable confidence. That era is already ending.
What is replacing it is what Abraham Sanieoff and others tracking this space have started calling the inference paradox. Cheaper intelligence encourages much greater consumption of intelligence. When individual AI tasks become inexpensive enough, organizations stop using AI selectively and start using it continuously. They embed it in more products, assign it longer and more complicated jobs, and run it in the background without waiting for a human to prompt it each time.
Agentic AI is the clearest illustration of this shift. When a human assigns an objective to an AI agent rather than asking a single question, the chain of computation that follows looks nothing like the old token-billing model. The agent plans an approach, retrieves relevant information, reasons through options, calls external tools, performs actions, checks its own results, and retries steps that did not work. One employee making what appears to be a single request may trigger an enormous amount of invisible computation behind the scenes.
Gartner has predicted that the inference cost of an individual agentic workflow could increase more than fivefold through 2028. That is not because the price per token is rising. It is because the number of tokens consumed per meaningful task is rising rapidly, and because the tasks themselves are growing more complex as companies discover what AI agents can actually accomplish.
Why Cost Per Token Is Becoming the Wrong Metric
Abraham Sanieoff believes that one of the most important mindset shifts businesses need to make right now involves abandoning cost per token as their primary measure of AI economics. In the early days of generative AI adoption, that metric made sense. Today, it is increasingly misleading.
A company could theoretically negotiate outstanding per-token pricing and still watch its AI costs spiral upward simply because it is running far more agentic workflows, embedding AI in far more customer touchpoints, and asking models to complete tasks that require dozens of internal steps instead of one. The unit price is only part of the story. Volume, complexity, and frequency are the other parts, and those three variables are all trending in the same direction.
The metrics that actually matter are shifting toward something more like the following:
- Cost per completed task, not cost per token generated
- Cost per customer interaction across a full journey, not cost per individual exchange
- Cost per automated workflow, measured end to end
- Return on investment per AI-assisted outcome in a specific business function
McKinsey expects inference, which refers to the actual running of trained AI models rather than the training process itself, to account for roughly 60 percent of AI compute demand by 2030. That compares to approximately 40 percent for training. This represents a profound shift in where money flows inside the AI ecosystem. The industry is moving from spending primarily on creating intelligence toward spending enormous and growing sums on using intelligence at scale.
When you frame it that way, the economic question for every business stops being about finding the cheapest model and starts being about generating the most value from each dollar of intelligence consumed. That is a fundamentally different optimization problem, and it requires a fundamentally different approach to AI strategy.
The Right Model for the Right Job - Why Architecture Is Becoming a Competitive Advantage
One of the developments Abraham Sanieoff finds most strategically interesting is the emergence of what might be called a model management hierarchy inside enterprise AI deployments. The idea is straightforward. Not every task needs the most powerful and expensive frontier model. Routing the right task to the right model could become one of the primary ways organizations control costs while still scaling their AI usage aggressively.
AWS has argued that organizations can achieve substantial inference savings by matching model size to task complexity rather than defaulting to premium models for every request. The logic mirrors how smart organizations already manage human talent. You do not assign your most senior and expensive people to tasks that a capable junior team member could handle equally well. The same principle is beginning to apply to AI.
A tiered architecture might look something like this:
- Small or locally-run models handling repetitive, high-volume, low-complexity tasks like classification, extraction, routing, and summarization
- Mid-tier models managing standard knowledge work, drafting, and analysis that requires reasonable capability but not the highest level of reasoning
- Frontier models reserved for difficult multi-step reasoning, high-stakes decisions, and complex creative or analytical work
- AI routers operating as traffic controllers, deciding in real time which model receives each incoming task based on its complexity and requirements
The competitive advantage in this landscape may eventually come less from having access to the best AI and more from routing millions of tasks to the cheapest AI capable of completing each one reliably. That is a capability that requires investment in architecture, tooling, and operational expertise. It is not simply a matter of picking a vendor.
Edge Computing and the Move Toward Local AI Execution
Abraham Sanieoff also wants to highlight a related trend that is beginning to reshape where AI inference physically happens. For the past several years, the dominant assumption has been that AI runs in large cloud data centers, and that accessing it means sending data to those centers and receiving results back. That assumption is being actively challenged.
Stanford researchers have argued for a future that is hybrid by design and local by default, where capable models running directly on devices handle the majority of appropriate work, while cloud models are called in for tasks that genuinely require greater computational power. Their research demonstrated approaches capable of recovering much of cloud-model performance while substantially reducing cost and latency.
This trend became particularly visible when Microsoft unveiled new hardware emphasizing local AI execution alongside cloud AI. The strategy allows workloads to be processed locally or remotely depending on factors such as task complexity, data privacy requirements, and cost considerations. The familiar debate between cloud and local is evolving into something closer to cloud plus local, with software automatically deciding in real time where each AI task should run.
For businesses, this matters because it introduces new levers for cost management and new opportunities to keep sensitive data from traveling to external infrastructure. It also means that thinking about AI infrastructure as purely a cloud expense may increasingly miss a significant portion of the picture. Edge and local execution are becoming legitimate components of enterprise AI strategy, not just edge cases.
What the History of Technology Tells Us About Where This Is Headed
Abraham Sanieoff believes that understanding this moment in AI history requires looking at patterns from previous technology waves. The dynamics unfolding in AI pricing and consumption are not entirely new. They resemble closely what happened with computing power, data storage, and internet bandwidth over the past several decades.
In each of those cases, the cost of the underlying unit fell dramatically over time. The cost of storing a gigabyte of data declined to near zero. The cost of a unit of computing power dropped by orders of magnitude. Bandwidth became cheap enough that streaming video became unremarkable. In none of those cases did dramatically lower unit prices cause society to spend less on the technology overall. Instead, lower prices made entirely new categories of application economically possible. Usage expanded to fill the available capacity and then kept expanding beyond it.
AI looks likely to follow the same trajectory. The falling cost of a token of intelligence is not producing a world where organizations spend less on AI. It is producing a world where organizations find vastly more things worth doing with AI. Every time a capability becomes cheap enough to deploy at scale, it unlocks a new layer of applications that was previously too expensive to consider.
Gartner's prediction that inference costs for a one-trillion-parameter model could fall more than 90 percent by 2030 is not a prediction that AI spending will fall. It is a prediction that the floor for what counts as a viable AI application will keep dropping, pulling more and more use cases into economic feasibility.
Abraham Sanieoff sees this as the central insight that businesses need to internalize right now. The next phase of AI adoption may not be defined solely by smarter models or cheaper tokens. It is likely to be defined by which organizations develop the operational sophistication to capture value from intelligence at scale, to route tasks intelligently across model tiers, to measure outcomes rather than outputs, and to build the infrastructure that makes AI a continuous engine of productivity rather than an occasional tool.
The question worth asking is not how much a model costs per million tokens. The question worth asking is how much economic value your organization can generate per dollar of intelligence. Companies that orient their AI strategy around that question are likely to be the ones that look back on this period as when they pulled decisively ahead. Abraham Sanieoff will continue covering the economic patterns, strategic decisions, and emerging architectures that define this next phase of AI adoption, so stay connected and revisit this space regularly as the landscape continues to shift.

Search
Recent Posts
Never Miss A Post!
Sign up for free and be the first to get notified about updates.
Newsletter
Stay In Touch
Featured Videos












