Beyond Intent: Architecting Networks that Think for the AI Era
We’ve been very energized by the market response since sharing our vision at the Open Compute Project a few weeks ago. This week, at SC’25 in St Louis, we take the next step: Aria switches will be running inside SCinet and will be on display in the Broadcom booth, where we’ll present in the theatre and unveil a few more aspects of our solution.
Network Automation Evolution
To frame the discussion around the Aria Networks solution, we must first recognize that the past decade’s model of network automation is no longer sufficient. The deterministic, rule-based approaches that were epitomized by Intent-Based Networking (IBN) fall short of today’s requirements. They are focused on state correctness and configuration verification, which are certainly useful but are now table-stakes. AI networking requires real-time performance optimization, adaptation and learning. Networks now need to sense, decide, and act continuously and in real time. The old model cannot get us there. This is the same shift that enabled self-driving systems to leap forward when they moved from rigid rule engines to probabilistic intelligence. Aria Networks was built for this new reality.
While AI Networking is the latest buzzword, many vendors have responded to the rise of AI networking by simply slapping a large language model on top of their existing networking stacks and claiming victory in their AI networking journey. At best, this approach provides limited insight, and at worst, can be quite dangerous. Indeed, if you were to ask AI to change your configurations, there is a non-zero likelihood it will make a mistake that can catastrophically bring down your network. It’s akin to putting a powerful language model in charge of aircraft navigation—it might sound smart, but without specialized flight controls, the outcome is catastrophic.
Specializing AI for Networking—The Foundation of Networks that Think
What we’ve learned from other domains that involve hardware and software interactions—like self-driving cars and robotics—is that it is critical to specialize AI for the domain. Specialization requires 3 components: optimized hardware (2) a reimagined software architecture, and (3) fine tuning of the AI models using data that’s native to the domain.
Let’s take self-driving cars as an example. The car itself needs to be able to collect telemetry across many dimensions using high-resolution cameras, radars, and, in some cases, lidars. The latest evolution rethinks the entire design of a car, taking out historical artifacts such as the steering wheel or the pedals and replacing disjoint systems with cohesive computers. The car software is fundamentally different, requiring an architecture that can ingest fine-grain telemetry at scale, and translate it to driving controls—accelerating, braking, steering, etc. Last but not least, the AI models powering the controls need to be trained on domain-specific data, in this case countless hours of driving videos across a wide range of scenarios.
The same methodology applies to the networking use case. Let’s describe how.
- Optimized Hardware: As with the self-driving car scenario, networking hardware needs to be optimized to collect fine-grain telemetry across many critical parameters. This requires a switch and NIC ASIC capable of collecting this high-resolution telemetry, and featuring the necessary on-switch compute, bandwidth, and memory to process and store the data in a sustainable fashion.
- Reimagined Software Architecture: The entire software stack needs to be reimagined in a way that can digest, store, and react to fine-grain telemetry at scale, in topologies supporting 100,000+ GPUs. State-centered approaches—optimal only for static configuration and deterministic outcomes—are not enough. The architecture must be telemetry-centric and inherently distributed, as actions can be taken at the ASIC level, the switch level, the POD level, or in the cloud. The architecture needs to be modular and be self-similar across these domains, with the difference being the resolution on these actions—from microsecond resolution and millisecond reaction times at the ASIC level, all the way to complex issue resolution in the cloud in seconds.
- Specialized AI Models: While AI models in networking can be relevant across all layers, they can look extremely different depending on their time scale and function. A model detecting load imbalance at the packet layer in one switch or one ASIC looks very different from a model that predicts that a transceiver is about to fail. And neither of the models have any resemblance to the LLMs that are at the core of the operator’s interaction with the network. And while in other domains, getting data is a challenge for newcomers, giving incumbents an advantage, our unique fine-grain telemetry (100-10,000x the existing solutions) will be instrumental in fine tuning these models very quickly.
Our experience over the last year has made one thing clear: building Networks that Think requires a new architecture and meticulous attention to detail. We invite you to join us in the months ahead as we share more specifics on these essential capabilities.
If you’re attending SC’25, we invite you to come see this in action. Our switches will be powering SCinet, and we will be hosting live demonstrations and discussions at the Broadcom booth. Follow us on LinkedIn and X to stay updated, and we look forward to continuing the conversation on the future of Networks that Think!