Optimize AI Traffic on the WAN with Ciena’s Integrated IP Networking (Sponsored)

Optimize AI Traffic on the WAN with Ciena’s Integrated IP Networking (Sponsored)

Heavy Networking

B2July 31, 202654 min
Play

As more applications and services rely on AI inferencing, that means more traffic on the WAN. Yes, bandwidth can solve a lot of problems, it’s not unlimited and AI inference traffic still has to share pipes with other flows. That means things like WAN latency, determinism, security, and traffic engineering are just as important as... Read more »

Transcript

Automatically generated from the audio. May contain errors.

Welcome to Heavy Networking, the podcast for listeners with strong opinions on EVPN VXLAN versus TradCore. I'm Drew Connery-Murray. Ethan Banks is away on a walkabout. Today, I'm talking optics and more with sponsor Sienna. And if you're guessing this is an episode about optics and AI data centers, not quite. Today's focus is the WAN infrastructure that carries

traffic among users, businesses, and the data centers running AI inferencing workloads. We're going to delve into the evolution of IP networks and coherent optical technologies that enable the movement of AI data across the internet. As more applications and services are relying on AI inferencing, that means more traffic on the WAN. And while more bandwidth can solve a lot of

problems, bandwidth is not unlimited, and AI inference traffic still has to share pipes with other flows. That means things like latency, determinism, security, and traffic engineering are just as important as the size of the pipe. So Sienna's here to talk about what it's doing with routing, switching, and coherent optics technologies to help engineers steer traffic

across the WAN as effectively and efficiently as possible. My Sienna guests are Rafael Francis, Senior Director of Product Line Management, and Vinny Santos, Director for Routing and Switching Portfolio Marketing. Rafael and Vinny, welcome to Heavy Networking. Rafael, I'll ask you to start by explaining the problem that AI inferencing presents for the WAN. What's different about

AI inferencing traffic that requires this kind of special attention? Yeah, thanks, Drew. And I I just want to say it's good to be here. It's my first time on the podcast, but I've certainly enjoyed and appreciated your content over the years. Thank you.

Good to be here. So, yeah, before I get into the details of inference, just to set a little bit of context. So, you know, Sienna, as a company, we offer solutions to service providers, hyperscalers, neoscalers, what we call neoscalers. Others refer to them as neocloud providers.

Uh-huh. And so these all play a role in the whole ecosystem behind AI-driven investments that are happening in networks. So, you know, we have kind of a front row seat and we have this, you know, distinctive view in that we operate at the intersection of IP routing and coherent optics and software control of all this. So they all kind of play a role in what's happening in AI investments.

And inference is really where a lot of the focus is now placed. Um, you know, to this point, there's been a lot of investment in, in, uh, the training of these AI models, the race to build the best models, you know, investing in, in AI data centers, GPU clusters in what we call a scale up, scale out and scale across parts of these networks. But really, it's this next, you know, those we refer to as the AI factories, if you will, right? The building of these LLMs. But this next wave or phase of investment is really, okay, how do you monetize all that investment, right? So a lot of money being spent, hundreds of billions of dollars annually, as you know.

And so now how do you monetize that? And that's where the consumption of these models comes into play from enterprises and consumers. And that's what AI inference is all about, right? How do we now interact with these models to gain insights and predictions and answers for things?

And obviously that inference for the end users, consumers, enterprises, has to traverse Metro and Last Mile networks to reach those end users. And that's where the service providers become a big part of that value chain. Um, and so, so that's, that's where things are now focused. Now, in terms of how, um, you know, AI inference is different, uh, certainly, you know, compared to just traditional broadband networks, uh, enterprise, residential, et cetera, consumers, certainly bandwidth is going to be increased, right? you know especially with multimodal AI you know more more video rich AI and imaging that's one way it's going to change so not just text-based AI not just chats I'm actually asking

my AI to build a video for me or give me some images for marketing or work or whatever exactly so bandwidth the beginning right what was that many LLM is just like the beginning of the story. We are like the tip of the iceberg. There's a lot that will come with AI and multimodal inference that will really change the model. And a lot of that is, some of that is also more

upstream, right? So now I'm providing a lot of context, which could also be, you know, multimodal. And so where traditional networks were very downstream oriented, I need some more upstream capacity perhaps now too, right? Latency is important, obviously, with inference. And network latency can matter more than compute latency, right? It's about that time to first token,

etc. So latency is going to, we'll talk more about that and some of the tools and things that we have for latency. And also always on, right? So now with Agentic AI, it's not just, you know, it's going to be 24 seven. It's not just us interacting with our applications, time of day sort of things. It's agents running all the time. So there's an expectation of always on.

So path and failure, domain diversity within the network, you know, sovereignty within the network. These are all, uh, you know, important inference in inference characteristics, if you will. But are these, you know, sort of net new problems or just an amplification of existing issues that network engineers have been dealing with? I think it's a fair question. I think you could

look at it as an amplification to some degree, right? The bandwidth that we talked about, the tighter latency requirements, the longer lived flows that AI, you know, that can result in, you know, from AI. The sovereignty, you know, issues I think were certainly prevalent within already cloud services and cloud infrastructure, right?

So I think to some degree it's an amplification, but the combination of all these things will translate to, I need to do something about my networks.

This is the opening of the episode. Open the player for the full interactive transcript with clickable words, translation and flashcards.

Open full transcript

Audio belongs to its publisher and is played from their feed. Rights holders can request removal — copyright & takedown policy

Episodes · Heavy Networking