The Rise of Inference
The last few months in the world of AI have been a stark contrast to where we were a few years ago. In an AI time far, far away (2024-25 in reality), all the talk, press, keynotes, etc., focused on training or the backend for those of you who are more infrastructure-oriented.
I went to three major AI events this past month. Two themes kept surfacing. First: to unlock improvements and economic benefits for customers, AI infrastructure requires a co-design approach that starts with hardware and extends to software, creating a cohesive full-stack experience. We identified this early on at Aria. This is why we built Deep Networking.
Second, inference was a massive growth area; as a result, it drove new use cases and, in turn, new solutions.
Previously, the prevailing belief was that inference wasn’t that complicated and that most inference workloads could comfortably be served by a single GPU server.
So what changed?
It comes down to two things arriving at once: unprecedented, exponential demand, and severe scarcity. You don’t have to look far for the evidence.
WARNING: If you suffer from Arithmophobia, you may want to tread carefully for the rest of this article.
Demand
Dell Inc. stated in its most recent earnings report that inference is past training and is pure demand, and that tokens driven by inference will grow 87-fold to 3,600 quadrillion tokens by 2030, while training demand will grow 5X by 2030.
If you, like me, are wondering what that looks like written out: 3,600,000,000,000,000,000.
The largest use isn’t by humans; it’s by agents. In the Q2 2026 Cloudflare earnings call, Matthew Prince, CEO, said that in five years, non-human traffic will be as much as 1,000 times as much as human traffic. Not because human traffic goes down, but because non-human traffic is growing so fast. Goldman Sachs is estimating nearly 120 quadrillion tokens (agent and non-agent) by 2030.

A year ago, Google served 480 trillion tokens a month. It now serves 3.2 quadrillion a month. A 7x increase in a year. The token count did not grow because we suddenly had 7x more humans using Google. It grew because each question takes more work to answer. NVIDIA’s Groq 3 LPX launch brief puts an agent at 15x the tokens of a normal chat request.
OpenRouter, which Stripe recently agreed to acquire for over $7 billion, routes traffic across more than 500 AI models and processes over 25 trillion tokens weekly, roughly five times what human users produce. The platform’s agentic volume has jumped 14x since February 2026.

Scarcity
All of those quadrillions of tokens have to run somewhere.
Scaling traffic is usually a matter of spinning up more resources. In AI inference, demand runs headfirst into the physical limits of silicon production, power grids, and datacenter buildouts.
The clearest place to see this bottleneck is the GPU rental market, where standard hardware economics have completely gone crazy.
The H100 shipped in 2022. Two generations of chips have come out since then. By the usual rules of hardware, the daily rate should have fallen. It hasn’t. SemiAnalysis has the H100 spot composite at $2.79 in July 2025 and $2.82 in April 2026, flat for most of a year while newer silicon shipped around it.
But wait, there are more numbers (I did warn you).
- Long-term contracts are up 40%: One-year H100 commitments climbed from $1.45 to $1.95 up to $2.10 to $2.70 per hour, while on-demand supply is largely sold out.
- The same chip carries a 5x price spread: On September 1, CCIR listed identical H100 compute from $2.15/hr (interruptible on a smaller cloud) to $10.53/hr.
- Legacy silicon is appreciating: The V100 is two generations behind the H100, yet it rents for more now than it did years ago. Lambda listed an eight-V100 instance at $0.55 per GPU-hour in May 2020; they list it at $0.79 now, and both configurations were out of stock last month.
The assumption was that every new generation would collapse the market for the one before it. That never happened.
Why?
Efficiency in AI does not free up capacity; it creates new demand. When OpenAI cut pricing for Terra and Luna in late July, usage exploded by 5.6x and 13.8x as previously uneconomic applications suddenly made sense.
Cheaper tokens don’t eliminate the need for GPUs; they ensure every generation of compute stays fully booked.
In the next blog, I will discuss the new designs this astonishing demand, combined with scarcity, is driving and the critical role the network and Aria play.
And I promise to try and use fewer numbers.