Holding Inference Together
In Part 2, we showed how GPU scarcity forced the disaggregation of prefill and decode, moving the model’s KV cache, its working memory, out of the GPU box and onto the network fabric.
As if we needed reminders of how fast this industry is moving, frontier teams are already splitting how models read context from how they think: carving out Feed-Forward Networks (FFN) and bouncing tokens across distributed Mixture-of-Experts (MoE) pools on the fly.
I promised fewer numbers in Part 2; I never promised fewer acronyms, but I’ll still spare you the five-paragraph explainer.
Once working memory and layer dispatch live on the wire, inference becomes a distributed computing problem with the network directly in the critical path.
That dynamic traffic hits production networks, where sizing in advance is no longer enough.
Tuning it by hand
Lossless Ethernet works. Priority Flow Control (PFC) and DCQCN run at scale in production right now. The challenge is about operating them.
PFC pauses the senders upstream of a congested port. Those pauses do not stay local. They spread across the cluster. DCQCN is worse in a quieter way: its settings have to match one fabric, one topology, one workload. Get it right in the morning, and it drifts by afternoon, because per Part 2, the traffic never stops changing.
Someone has to correct that drift. Historically, that someone is a specialist, doing it by hand. This is why RoCE earned its reputation, and why plenty of operators paid the premium for a proprietary stack rather than live with it.
Also, a human adjusting static thresholds cannot track a workload that reshapes itself every millisecond, and no amount of expertise fixes that, because the mismatch is in the clock speed, not the skill.
Which points to what replaces it.
More on that later. Yes, that is a teaser.
What it costs
Ask anyone building an AI Factory what good looks like, and you get a different answer depending on who you ask.
The agent team wants Time-to-First-Token and steady inter-token latency. The batch team wants throughput and Token Hardware Utilization. Finance wants cost per million tokens and tokens per megawatt. All three are correct. None of them can win at once.
That is the Pareto frontier. Optimize for latency and utilization falls. Optimize for throughput and queues build until the agent feels dead. Picking a spot on that curve is strategy, and it is a legitimate thing to debate about in meetings; we have.
Fine, it is a tradeoff; every system has tradeoffs. Tune it and move on, one might say. You can, if the workload stays put. But static tuning only works when traffic patterns are fixed.
In production, the target is constantly moving. An agent loops through dozens of tool steps, suddenly dumps a 100,000-token document into context, and fan-out traffic collides on the same switches.
The moment a buffer fills, even for a few microseconds, GPUs stall waiting on packets. For a human reading a single chat stream, a brief pause is annoying. But for an autonomous agent running a 50-step recursive chain, those microsecond stalls compound into multi-second hangs, blown latency SLAs, and broken timeouts. You lose on latency and utilization at the same time.
What we build
That is what Deep Networking is for.
It starts with telemetry at 100 to 10,000 times the resolution of legacy tools, from the host through the fabric. Per-second sampling doesn’t measure them badly; it doesn’t see them at all. And those microsecond events are what decide whether a KV cache transfer or an MoE dispatch stayed hidden behind compute or stalled the cluster.
Resolution alone is not enough; you need correlation: network data and host data on one timeline.
In our lab, Aria noticed an expensive GPU node suddenly running unusually cool. Most people would see that and assume producing tokens at a lower temperature was good, or that the hardware was faulty. Aria correlated host events with the fabric and discovered the GPU was cooler because of network congestion.
An expensive accelerator sat idle, waiting for the network.
When a request comes back slow, you can finally see what caused it: the fabric buffer, a degrading optical transceiver, the NIC, the host OS, or the model server. An Aria network is engineered so the network isn’t the bottleneck keeping you from your Pareto frontier.
The self-tuning future
In Part 1, we showed that the token explosion is overwhelmingly driven by autonomous agents. In Part 2, we showed that hardware scarcity forced model memory onto the wire.
As inference architectures fracture into distributed attention, FFN, and expert pools, static networks cannot survive. Sizing switches in advance is no longer enough. You have to see the fabric at the microsecond resolution where failures actually occur.
Seeing is half of it. The other half is what happens next, and whether anything is remembered.
Go back to the specialist. A specialist tunes the fabric in the morning. By afternoon the settings are wrong, and they do it again, from scratch.
A control loop does not start from scratch. The loop keeps what it learns. A congestion event it clears, a transceiver it catches going bad, a traffic shape it has not seen before: all of that becomes history, and next time it tunes against the history instead of starting cold.
And the history keeps growing. As Part 1 showed, cheaper tokens do not mean less demand. People buy more, traffic goes up, and the loop gets more to learn from. It’s the flywheel concept I discussed previously.
AI broke traditional networking. What replaces it is a Network That Thinks: an agentic network running self-tuning control loops in machine time, adapting as fast as the software on top of it.