NetworkANGLE
NetworkANGLE is focused on the evolving role of enterprise networking as a strategic foundation for digital transformation and AI-driven operations.
The blog provides objective, research-driven analysis across campus, data center, WAN, cloud, and edge networking domains, with particular emphasis on how modern architectures must adapt to support AI workloads, distributed applications, sovereign requirements, and real-time performance expectations. NetworkANGLE examines architectural decisions, operational tradeoffs, and ecosystem dynamics shaping next-generation networks.
Typical topics include AI-ready networking, Neo Clouds and GPU-as-a-Service infrastructure, self-driving networks, fabric architectures, the evolution of SD-WAN and NaaS, and the convergence of networking, security, and observability. Content is written for execs, infrastructure leaders, and network architects who need to translate technical capabilities into business-relevant outcomes such as agility, resilience, scalability, and governance.
In short, NetworkANGLE frames the network not as plumbing, but as a programmable, intelligent platform that increasingly determines enterprise and service-provider competitiveness.
NetworkANGLE
Gilad Shainer, NVIDIA
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Gilad Shainer, SVP of Networking at NVIDIA joins theCUBE research hosts Bob Laliberte and Dave Vellante for a NetworkANGLE
In a few short years, artificial intelligence has changed how organizations think about their technical, their business, and their operating models. Firms are moving beyond experimentation into production and tapping into manufactured intelligence, which is coming out of AI factories. Now, these factories are not only capable of training and serving highly sophisticated models, but they're increasingly supporting day-to-day operations with real-time inference, powering agents that reason, they make decisions, and they take actions. As this wave of AI advances, one thing has become abundantly clear. Networking is not about a bunch of pipes that connect infrastructure. Rather, the network has become a fundamental component of the AI computing platform itself. Now, at the same time, the industry is seeing a significant shift toward Ethernet-based AI networking. Questions about openness, standards, scalability, performance, and operational simplicity are becoming central to enterprise AI strategies. Hello and welcome to Inside the AI Factory, how networking fabrics power agentic AI at scale. This is a special editorial deep dive brought to you by SiliconANGLE and theCUBE Research. My name is Dave Vellante, and I'm here with my co-host, Bob La Liberte, who is one of the top networking analysts in the industry. And today we're joined by Gilad Chener, who is Senior Vice President of Networking at NVIDIA. NVIDIA, of course, has been instrumental in driving the transformation of networking to support accelerated computing. Welcome, Gilad. It's great to have you back in theCUBE. Thanks for making some time for us.
SPEAKER_02Yeah, happy to be here.
SPEAKER_01So listen, Bob and I have been spending a lot of time talking with enterprises about networking for AI, and we want to explore how networking is evolving in the AI era, why Ethernet is gaining momentum. What organizations should think about as they build next generation of AI infrastructure. And Bob, you've recently completed a pretty deep study on networking for AI research. Why don't you kick us off with some of the findings there?
SPEAKER_00Yeah, absolutely, Dave. Thanks. I mean, I think one of the points you made in your opening was really clear. For the longest time, organizations have considered networking just the basic plumbing. Yet what our research shows is that's really changing. And one of the strongest findings from our recent research report, which was the networking for AI research, we were talking about, is that network has become a strategic business priority. And it's not just like half the organizations think that, but nearly 95% of organizations told us the network is more important to meeting business goals than it was just two years ago. So that's a huge change that we're seeing. At the same time, as you mentioned, we're seeing AI evolve from these isolated training jobs into highly distributed inferencing and increasingly agentic workflows where the AI agents are continuously communicating with one another. So our research shows that there's a lot of growing concern about the impact that these agentic workloads will have on the enterprise workloads. So to get things started, Galat, I wanted to turn to you. And, you know, based on that as the background, how do you see agentic AI fundamentally changing the network requirements inside an AI factory compared to perhaps the first generation of AI infrastructure?
SPEAKER_02Yeah, um, it's a good it's a good question. And actually we can go back a little bit in time uh and see how all started. Um AI is a distributed computing workload, which means it's a workload that needs to run on a collection of compute engines, and only when you have the collections of compute engines that can work as a single unit, you can actually run AI, you can actually run distributed computing workloads. Now, in the beginning, the first workload were training workloads, and training workloads uh seemed hard at the time because it did require a completely change in the way that data centers were built. If for other workloads, the traditional enterprise workloads, where the only thing that you need to do is to access a compute engine, run your workload on a single compute engine and get the result back, when you're running training or pre-training, you actually need to access multiple compute engines, and you need all of those compute engines to work as a single unit because you need to bring the power of many compute engines together. And it was hard because now the network has become a key element. The way that you connect compute engines will mandate how that infrastructure, how that data center, how that factory is going to work. If you have a network that cannot enable the collection of compute engines to work together, then you just have a server phone. If you want to build an AI supercomputer to run training, you need to bring a network that can send data between GPUs, for example, at the same time be able to move the GPUs or transfer data between the GPUs in a full synchronization mode and so forth. Those are high-performance computing networks. This is where we started. And and people gave uh or or knew that training is hard, but at the same time, the thought that inferencing would be easier. Um actually it's it's wrong. Okay, and we see it today. Inferencing is very, very complicated. So training was complicated. Inferencing is much harder than that. If you look on inferencing today, um and and where we are with agentic AI, the AI factories are not just GPUs and network. Now we have many compute engines and different compute engines that need to work together. We have GPUs, we have CPUs that previously were just supporting GPUs, but now they are a key element. They're a key element in running inferencing, um, in in generating tokens. So we have GPUs as as an important compute engine. We have CPUs that have become very important engines, and then we have LPUs that are also very important compute engine. So now we have different compute engines, and and those engines need to be connected together, those engines need to work as a single unit, you need to move um uh element of AI workloads from GPUs to CPUs to LPUs and back to GPUs and CPUs and so forth. That then actually it's much more complicated. Okay, much more complicated. And in order to support that, we needed to develop more infrastructures or new infrastructures in order to build those AI factories that can generate tokens and tokens actually generate money for enterprises. So if in the past we were focusing on East-West network, a scale-out network, and scale up was mainly uh GPUs, for example, so within the box. Now we're talking about scale up network that goes and connect tens and hundreds of GPUs. We're talking about scale-out network that connect hundreds of thousands of GPUs and CPUs and LPUs. We're talking about new storage infrastructure that supports KV cache and contacts and do that in a very efficient way. We're talking about access network that enables secure access of many users that are gonna run their workloads. And we're talking about scale across network that needs to connect multiple AI factories together because we need to harvest the compute capabilities of many data centers in order to do the next generation of AI workloads and so forth. So we grew to have five different networking infrastructures, each of them is purposely built for a specific mission in order to essentially take all of those compute engines, the different compute engines, and make them work as a single unit.
SPEAKER_01So, Gilad, if I could follow up on that. I mean, what you described in going back to the beginning and when training was sort of the main workload, you described an architecture which some would describe as monolithic, and and they maybe use that term as a pejorative, but but my inference, no pun intended, in listening to you is that approach, your architecture allows synchronous activity across all the whether it's GPU or uh LPUs or CPUs, in a synchronous approach, sharing memory. Um this this ties into you describing uh sometimes it's not just a one-chip problem, which is sort of summarized. So my question is when you designed the original network, how did you have to evolve that to support this heterogeneity of compute and still enable that, if I get it right, that synchronous connection? I mean, I hear others talking about chiplet architectures, which I infer is maybe not synchronous. I wonder if you could sort of cut through the noise there and help us understand architecturally how you've evolved the network to accommodate that.
SPEAKER_02Yeah, um, it's it's a good question. And you're right. Um, if you look on um NAI pod a few years back, you would see that there were GPU recs, which were connected with NVLink, and it was NVLink 8, for example, a few years back. Um, and then we had a uh a scale-out network that connects those racks together, and pretty much that was the pod. Um today the pod looks completely different. You know, we have the GPU racks, and we have a large scale up domain that connects tens and moving to hundreds of GPUs. We we we have the scale-out domain that grew in size in size, um, and and we have those racks for Spectrum X. We have new racks for CPUs, just CPUs, because they're becoming an important factor as you mentioned. It's not just the GPUs. We have different racks for LPUs, we have new racks for storage, for holding context, and there's also rack for uh for enabling the north-south access into the pod. The pod has become huge, and different compute engines, different infrastructure racks, for example, and so forth. Now, the way the way that you need to look on that from a design perspective is that entire machine, uh the GPUs, the CPUs, the LPUs, storage needs to work as a single unit. And this is where we're describing uh an extreme co-design. You need to bring elements that are purposely built to serve a connectivity that is needed for those the specific compute engines and then how they interact with the other compute engines that exist. A scale-up network is actually something that needs to be purposely built to deliver massive amount of bandwidth, extremely low latency, very high uh message rate or packet rate on the network, and in network computing engines. Because the purpose of a scale-up connectivity, the purpose of scale up network is to take compute engines and to make them one. So when we look on uh NV NVL72, it means that those 72 GPUs are actually behaving as a one GPU, as a REX scale GPU. And when we go into NVLink 576, for example, then you have 576 GPUs that behave like one GPU. Therefore, the scale-up network is very complicated and it's to be purposely built together with the GPU in order to build that REC scale GPU or that 576 scale GPUs, for example. Now, a scale-out network needs to connect multiple compute engines now together, not just the scaling out of GPUs, but also to connect the LPUs to the GPUs and also to connect the CPUs between themselves and the CPU to the GPUs. And here we have different flavors for example, for example. We have the scale out network that brings less amount of bandwidth, but it needs to be very synchronized. Latency is important. Uh j tier is critical. J tier is critical, and the reason that JTR you you want to eliminate JTIR is because that pod needs to work as a single unit, which means you want to make sure that all the computational activities are completely synchronized. You want to make sure that every GPU, every CPU, every LPU get the data at the same time so they can compute at the same time and then exchange data at the same time and do it over and over and over again. If you don't have a purpose-built network, if you take an off-the-shelf network, they didn't really care about synchronizations for uh distributed computing workloads, that network may create delays. And delays means that the compute engines will not be synchronized here, and one will need to wait for another. And when one waits for another, essentially your AF factory becomes a server form. Okay, and actually, you are not able to generate tokens. So that's why there is a heavy investment in making sure that everything is synchronized. And then for LPUs connected to GPUs, we actually created a low latency version of Spectrum X because we wanted to have a shorter latency between what works on LPUs and then work what what works on the GPU, so they can actually move data between them very quickly. And then we brought storage and context. And here we build a purpose uh a network with Bluefield, with Bluefield 4, the ability actually to run the storage operations on a Bluefield, on a DPU, on a storage controller, which sits on a GPU box on one side and sits on a storage box on the other side, and make that very effective uh from from sending data to the storage and then bring the data back to the storage and so forth. So this is where we actually took Bluefield 4 and optimized BlueFit 4 and from a DPU that was handling access network, now it's actually running the storage infrastructure, the storage framework, and it's the best storage processor and storage control that we can we can put our hands on it. Um and then obviously there is more and more work on Scale Across and so forth. Every infrastructure serves a specific need. Every infrastructure is purposely built, and it's built in a in a in a in a in a extreme co-design approach. So the way that the GP works is completely synchronous, that the way the connectivity works, and then the software that runs on top of that can leverage everything that exists in every aspect of that. That's an amazing engine that is being built. And this is why it's so complicated, and this is why it's so complex, and this is why you need to make sure that everything works with any other element in that AF factory.
SPEAKER_01Yeah, so we want to come back, by the way. We'll we'll do this later on in the in the discussion and talk about what that all means for economics. But we want to dive into uh the Ethernet discussion. When when you guys announced Spectrum X, uh Gilad, you and I talked about the scale up, scale out, scale across. Um, and you've heard the criticisms uh uh around the whole openness conversation. Uh people say, well, Spectrum X, you described it as you had to create a specialized version of Ethernet to accommodate AI. Some say that Spectrum X is a proprietary implementation uh of Ethernet. You've emphasized that you use open standards like like Rocky. You help us separate perception from reality. What makes Spectrum X an open Ethernet architecture? How do you respond to the criticism?
SPEAKER_02Yeah, first, Ethernet it's it's open. You know, Ethernet is an open architecture by definition. That's that's number one. Um, number two is that we're not using any proprietary protocols in Ethernet. The way that Spectrum X works, it's actually using the standard Ethernet protocols. You can connect Spectrum X switch to any other Ethernet switch and it's gonna work, and you can connect um our uh Ethernet NIC to any other switch, you connect our switch to any other NIC, and it's fully interoperable because it's running Ethernet. Um, even for there, we uh fully enable any operating system to run on our switches. So we were um uh one of the major contributors, but for example, to SI, the the switch up structure layer, which is a fully open uh uh interface. We contribute a lot to Sonic, which is an open source um network operating system that exists, and we have many customers that are actually using that. Um we collaborate with Cisco and Cisco running their network operating systems on our Switch ASICs. Uh we work with Meta, and Meta has their FBOS, which is their network operating system, runs on Spectrum Switches. Anyone, any any of our customers can design any network operating system and run it on our switches. So there is a fully openness from the software perspective. It's it's running the the network, uh, the internet network protocols between switches and so forth. So everything is using starred elements. The the magic of Spectrum X, okay, the reason that Spectrum X was purposely designed for AI is in the way that we implement the activity or the way they implement what the switch does versus what the Nick does. Um, remember when we when we talk about Ethernet, there is no one Ethernet switch out there. There are many kinds of Ethernet switches. Even before we started with Spectrum X, you know, even before we started with building a purpose-built Ethernet for AI, there were different kinds of Ethernet already in the market. Different kinds of internet switches already in the market. There were switches that were designed for enterprise, uh uh hype uh uh virtualized enterprise data centers. Those were kind of switches that has um less number of ports, a lot of functions on them to support virtualization and so forth. And that was one kind of switches. There was a second kind of switches, or there still is a second kind of switches that were built for hyperscale clouds. Those switches include um a lot of ports, much more, higher number of ports, less functions, and they are using for uh or used for uh um supporting large infrastructures and single server workloads, for example. There was a third kind, still is there is a third kind of Ethernet that was built for telco and and DCI and long distance, which the implementations of those internet switches are based on debuffers, so they can run or support long distance for um off the shelf or traditional workloads. Before Spectrum X, there was already multiple kinds of Ethernet, multiple kinds of switches, multiple kinds of implementations of protocol under Ethernet in that sense. Spectrum X continue to use the the Ethernet protocol that are there, that are set, that are uh open as part of the internet, uh uh as part of Ethernet. But the way that we built the infrastructure is was to support distributed computing workloads. So if previous or other other other uh um uh architectures of internet switches uh were mainly focused on let's try to put everything on the switch, okay? Let's do the full implementation of the switch. The switch will uh choose the routing schemes based on Ethernet protocols, and the switch will make sure that the data uh is is being sent or or received in order and transfer in order all the way to the other side and so forth. And that those were the implementations that existed in the other Ethernet or the different uh of the shelf Ethernet that exist, and they were built in a way to support the the workloads that they were designed for. Okay, the reason that we have uh uh switches with debuffers is to support workloads that runs on long distance or traditional workloads that runs on on long distance. There were different implementations of Ethernet versus switches that were built for hyperscale cloud, single server workloads where they have much smaller, smaller buffers and more ports and so forth. So with Spectrum X, we we we we we designed the infrastructure in the way that can actually support DCB computing workloads, and that's the magic of Spectrum X. That's why Spectrum X was purposely built for AI. And with Spectrum X, we knew that the number one problem that we need to solve is JTR. Because all other Ethernet options that exist in the market, all the off-shelf internet switch options that exist in the market, didn't really care about G tier because GTA was not a problem that they were trying to solve. Now, in AI, you have to solve GTR. If you don't solve GTA, your infrastructure, your AF factory will become a server farmer, not an A-factory. And in order to solve GTR, in order to eliminate G tier, we needed to spread the functions, spread the networking functions, the networking implementation between the switches and the supernix. Because the switch uh cannot uh do uh the best data distribution, the switch cannot cannot freely distribute traffic, and make sure that you leverage all the network uh paths that exist, uh, and at the same time uh make sure that the data is being sent in order to be consumed by the receiver. It it it's impossible. So we needed to split the functions, and the switch is tasked to do unconditionally uh traffic distribution. So all the traffic will use all the paths that. Exist and then we eliminate JTR across the entire infrastructure. And the switch does it with getting data and understanding what's the situation of other switches around him. This is the IP, this is the implementation of NVIDIA. This is the implementation of starring internet protocols. On the other side, we task the supernik to control injection rate and then put the data back in order in the GPU memory. So essentially the GPU can actually work on the data in the right order. That's why Spectrum X is an infrastructure. It runs started internet protocols. You can run any operating systems on top of that. But the implementation on how things work, that's what makes Spectrum X a purpose built for AI.
SPEAKER_00And Galah, the interesting thing here is that when we talk about standards and Ethernet and ultra-ethernet consortiums and things like that to help drive the standards, the reality is that NVIDIA is moving so fast that you have to innovate. We talked about that extreme co-design and so forth and making sure the performance is there. So you're going to be always on that leading edge of innovation and then feeding the standards and driving that technology into the standards. Is that is that a fair statement?
SPEAKER_02So, first, you know, we we are part of a lot of uh uh standardization uh organizations and we continue to enhance the specifications and continue to enhance standardization for the standardizations. We are um uh uh working with the IBTA consortium, for example, which is where Rocky is being standardized and continue to enhance Rocky capabilities and so forth. We are part of UEC, we're part of um Isang, we're part of many consortiums, and we actually contribute to that. We continue to contribute a lot of um open code, open source code to Sonic to continue to enhance the open source network operating systems of the switches and so forth. So we continue to help and enhance and provide the capabilities to uh the different consortiums and driving more and more elements into sterilizations. At the same time, we need to continue and work very, very fast because um we are in a in a in an era where every year there is a new generation that is being released. In in the past, if you look on the cadence between generations, it was around three to four years. So you can work on a generation and then have the generation being used for a few years, and you have several years to work on the second, next generation, and so forth. Um, nowadays it's it's every year. So every year there is a new release of a new platform, new GPUs, new CPUs, new new LPUs, new switches, new new software and new software elements, new accelerations, there is new workload, and so forth. The cadence is much faster. So obviously, uh we need to continue and and innovate and provide more capabilities and in order to support larger um and more diversified from computing engines, um, AI infrastructures, air factories. And at the same time, we continue to uh uh work with the ecosystem and drive more standardizations and more element and so forth.
SPEAKER_00Yeah, no, I think that's that's great. I think one of the, you know, you brought up the extreme co-design we talked about a little bit, and you know, essentially that ability to optimize the GPUs and networking storage, software, frameworks all together as a complete system and the importance of that. And I think what I was hoping that maybe we could also cover for everyone who's listening is how important it is to have the networking be in that co-design, right? It's part of that philosophy and how it's become, I think you've touched upon some of these already, but how it's become such an important determinant of the overall AI factory performance. And you know, as organizations are thinking about building out their AI factories, how important it is to have the network as part of that extreme co-design when they're deploying an AI factory.
SPEAKER_02Yeah, I I think I think the the best way is uh I can give you an example. I can give you an example that shows why co-design is so important. And um I'll I'll take I'll take NVLink as as one example. We can take examples on all the infrastructures, but I'll take I'll take NVLink or scale up as an example. Um and I can start with maybe something um um something simple to illustrate um what what we do and why it's so important. Um and I sometimes I use uh an example of pizza, baking pizzas. Okay, let's say that you are an owner of a pizzeria and you're baking pizzas and um you have the entire process of making the dough and putting it all together and put that pizza in the oven and and wait for the pizza to be baked, and then you move the pizza into the uh delivery person, delivery person will you know bring it to bring it to the consumer. Now let's say that um your mission is to accelerate the time of preparing pizzas. Um let's see what what can you do at that point. Um one way is maybe to try and and bring more workers or bring more ovens. Um, if you do that, you can actually have more pizzas being created at the same time, but the time per pizza stays the same. So you're not really helping or not reducing the time from getting the order and having the pizza delivered. Um, another approach is looking on the entire process, and I says, Well, I don't have to have to to run the entire process in one place. What if I'll put the oven in the car? So I actually can prepare the pizza, and then when the pizza is being delivered, I can I can bake it. So in that case, you can actually reduce time. And this is a good example of co-design because you had two different things. You have a compute engines on one side, you have a networking on the other side, and once you start having a co-design between them, you could accelerate things. Now, the same the same example valid for NVLink. I might take that as an example, and this is where I'm referring to in network computing. The same approach. In scale up, we're taking a lot of GPUs, for example, and we want them to work as one unit. And one unit means that obviously there is a massive amount of bounds that needs to be go between them, but remember that every GPU computes something, and then every GPU needs to communicate that result with every other GPU, and then all other GPUs will combine those results, and then they can do the next phase of computational phase and so forth. Now, what we do here is instead of having the GPU send the data to every other GPU, and then every GPU needs to calculate the answer by itself, we take that information, send it to the switch, the N-Vinning switch, the NVINE switch would actually do the computational element of combining those results from all the GPUs and then send the combined result back to the GPUs. So the network is actually active in computing uh workloads. Now, the algorithms that need to be run on the switch as part of that computational phase is needs to be based on what algorithms run on the GPU. So it's not just co-designed between the GPU and the network where the compute part is happening in both and needs to be synchronized, it's also co-designed with the software that runs on top of that because the algorithm of the software needs to be implemented in the way that you can run part of that on a GPU and part of that on the on the on the switch, and from generation to generation, those algorithms are being modified and optimized, and the four the in-network computing elements need to continue to be uh updated on the switches and the GPUs and so forth. This is an example of an extreme co-design. It covers the software, the compute, and the network together.
SPEAKER_01So if I take your pizza analogy a little further, the you have to make sure that those ovens that are in the delivery truck, uh the pizza that comes out of those are as delicious as the one that my chef makes back in the in the factory, if you will. Um, and I assume that extreme co-design uh allows you or it forces you, uh it brings other challenges that you have to address. That and I want to tie it into something you said earlier. A couple of times you said you then you just have a server farm versus an AI factory. And um and I'm I'm curious as to what the difference is. And if I use your pizza analogy, that you're dealing with those challenges, the way you deal with them is extreme co-design. You maybe have to make new inventions or new algorithms to make sure the pizza is as delicious. But what is the difference between an AI factory and a server farm?
SPEAKER_02Well, a server farm is a collection of servers that support single server workloads in a sense that every workload or the the compute capacity that the workload needs to consume sits within a server. It could be a single core, it could be single CPU, it could be a single server, but that's it. There is there is no need to have connectivity between the servers. Essentially you need one connectivity to get into the server and out of the server. And this is the way the traditional hyperscale data centers were built. There was one main network, an access network, get you into the server and out of the server, and that's it. Um when you when you deal with distributed computing workloads, it's not that you need to get into the server, you don't need to get move your data into the server and get the result back. The data needs to go to many servers, and then all those servers needs to work together because your your problem cannot be dealt with a single server. Your workload is too big to be dealt with a single compute engine. Your workload requires collections of many, many compute engines that needs to work as a single unit. So now you have a problem that needs to be broken into small pieces. Every piece needs to be sent to a different CPU, a different GPU, a different server. Every server will compute that piece and then will exchange the result with any other server, and then you move to the next phase of computational, do another round, and then exchange data again, and then do another round of compute, exchange data again, and so forth until you conclude that workload, and then you can actually send the result back. Okay, so if a server farm is just a collection of servers that each of them work independently with no indications or no knowledge of any other or no interaction of any other server, when you build an AF factory, all those servers, all those compute engines needs to work completely together and needs to be fully synchronized in order to fully utilize the effectiveness of that collection of compute engines and to maximize the performance and efficiency of the workloads that are running on top of that AF factory.
SPEAKER_01So if I try to put a oven on the AI or the server farm truck, I would have a truck that has dough, another truck that has sauce, another truck that has cheese, another truck that has pepperoni, and I have to put them together at the end, and it would not be as delicious.
SPEAKER_02I'll I'll I'll put I'll put the example as the the pizza is so big, okay. Here's here's a good example. The pizza is so big, okay? The workload is so complicated, that pizza cannot fit a single oven. Okay, no way it's gonna fit a single oven. So you need to break the pizza into small pieces, and each piece is gonna go into a different oven, and then you need the ovens to be so synchronized that they will finish at the same time. So every slice of that pizza will be the same heat, the same readiness, and and that there is there couldn't be a slice that is earlier and slice that is later, because then one bite's gonna be cold, another one's gonna be hot. Okay, and and once you have all those pieces, all those ovens uh work in a fully synchronized mode, then you can build a pizza that will be as delicious as it was just a single small one. I love it. We're Rob, we're bringing this is a much better example that this you know that demonstrate what does it mean.
SPEAKER_01It's beautiful. We're bringing you know AI to the masses, explaining networking with the terms that people can everybody can understand.
SPEAKER_00I love it. Well, I think who knew that pizza was so important to making having networking for AI factories. So I think it no, I think it's a great analogy. I think it really works. And you know, Glad you you've talked about networking right clearly as an integral part of the AI system, not just an interconnect, right? That's become really clear. And I just wanted to touch upon another area where you guys have really been pioneering, and that's the scale-up networking with NV Link. So I'm wondering if you could share with our listeners, you know, really what makes NV Link ideal for that scale-up networking?
SPEAKER_02Yeah, um NVLink, NVLink, it's uh it's uh it's actually an extension of the GPM, if you think about it. It's actually an extension of the GPM. Um if if a scale out network is gonna connect a lot of recs, you know, hundreds of thousands of GPUs, and needs to provide the low latency and the zero jitter and and so forth between them. Um, but they're still connecting a lot of GPUs together in that sense. Scale up network or NVLink network. That mission, that scale up, that NVLink mission is to create a GPU. It's not to collect is not to connect like scale out hundreds of thousands of GPUs and make sure they are fully synchronized and and so forth. Scale up mission is to create that GPU. So you need to create a single GPU unit, and that GPU unit is being built with multiple GPU ASICs. Okay, so now you have many ASICs, and then you need those ASICs to become one. And therefore, the amount of bandwidth, the amount of data that you need to support between those GPU ASIC is an order of magnitude higher than the data that you need to run or support on a scale-out network. Okay, so we're talking about order of magnitude, higher bandwidth that NVILIC needs to support. Second, it needs to be much lower latency, much, much lower latency. Remember that we are we are having GPUs, they need to become one unit, and therefore the communication between them needs to be very, very quick. Very, very quick. So the latency, the latency of NVILIC is much lower, much smaller. Um, or the communications are much faster than what we need from a scale-out perspective. So we're talking about order of magnitude of bandwidth, we're talking about almost order of magnitude of latency, and we're talking about order of magnitude, for example, on message rate, again from the same reason. And that's not enough because we also brought the e-network computing into NV Link. We also needed to bring Sharp into NV Link. We needed the NV Link to be part of that computing process, and that's from the same reasons of the pizza. You need to get things to work much faster, much faster, and it's not adding more GPUs will help me there. I w I need to move some of the computing engines to actually be part of NVLink. So NV Link is a computational network in that sense, it's an extension of the GPU, order of magnitude of bandwidth, order of magnitude of lower latency, order of magnitude of packet rate, and the sophisticated e-network computing. And that's what makes NVLink the best scale-up network you can find. And because it's not easy to do, and because it's the best scale-up network that exists, purposely built for scale-up AI, we we wanted to make sure that everyone can enjoy it. And obviously, of course, if you're using NVIDIA GPUs, you can you're gonna enjoy NVLink uh for sure. Uh, but if you build your own accelerators, if you build your own XPUs, you build your own SP uh CPUs and so forth, why won't you want, why don't you want to enjoy the capabilities of NVLink, the extreme bandwidth, the low latency, the message rate, the the the reductions uh in it for computing. And we when and and for that reason we created fusion. And with NVLink Fusion, people that our customers that have their XPUs design because they they wanted to build a purpose-build compute engines for their workload, for example, or they build their own CPUs, can actually take our scale up network and connect that with their XPUs, and with that be able to leverage everything else that Nvidia is being built. Everything else that Nvidia has created. So they have their XPUs, and now they can leverage the capabilities that we brought with NVILLink uh under NVILINK Fusion, so they can actually have a scale up between those XPU and accelerators, and and they can also use the racks that we built because they can actually put everything in the rack, and then they can enjoy the liquid cooling and the 45 degrees of liquid cooling all around that, and the copper backplane that we built, for example, and they can also take scale out network if they want to, and they can leverage Spectrum X and so forth, and the storage access and everything. They can leverage it, they can leverage all the infrastructure that we have created so they can focus on their XPU, release those XPUs, and then the entire infrastructure that will connect their XPUs to the rest of the AI elements, the rest of the AI factory that they built. They can leverage everything that Nvidia has created. Um, we see a lot of excitement around that because it saves time and actually people can enjoy from the greatness of the infrastructure that we build, and they can and they can bring their own XPUs, their own accelerator into their usage. So that's I think actually uh as well a great match.
SPEAKER_00Yeah, I'm glad you brought up NV Link Fusion because I was I think you launched it about a year ago, and it's really a great example of the ecosystem you're building and allowing other builders, enterprises to be able to leverage that technology for specific applications that they might have.
SPEAKER_02Yeah, that's another example of by the way, the openness of Nvidia, because you can look on the ecosystem of of NVLink Fusion, it's it's amazing, it's it's big and it's growing and so forth. Everything we build, we provide for our customers to use, and they can take every piece of that, they can configure every piece of that, they can run their own network operating systems on the switches, they can use NVLink Fusion and connect to the XPUs. The interfaces are completely standard, and um everything we do is is a very open infrastructure, and we have customers that are really happy to use it's a whole as a PC and and build their their optimized AI factories for their own purpose.
SPEAKER_01Yeah, and fusion recognizes that the world is heterogeneous. I'm sure you'd love for everybody to always buy NVIDIA, but that's not reality. And by by uh launching Fusion, you just you make the market bigger, I mean very clearly. You know, I want to uh come back to sort of economics. I mean, networking uh uh the discussion's always been kind of geeky, standwith and ports and fiber optic connectors and where to use copper and a very technical uh uh environment. But you know, enterprises today, they're talking about things like GPU utilization. There's a lot of talk about you know token maxing. We saw Alex Carp rant on CNBC the other day, uh which sets up kind of an interesting you know discussion. Uh, but it's networking has become, as we've described here, part of the system. It's not just this isolated set of switches. Uh so it's tied to business outcomes. So at the end of the day, you've talked about this annual cadence. We've seen the cost per tokens go down 90% a year, and they're they're gonna you're gonna continue to drive that. You have to, because right now people are you know complaining about about cost, you're gonna see a lot of cost optimization. But ultimately, you know, these are gonna be intelligence is gonna be very inexpensive. So, what role does networking play there? How should organizations think about networking as not only a cost decelerator, but a business accelerator?
SPEAKER_02Well, the the you you you said it correctly. You know, that it's it's the most important thing is the cost per token. And I think NVIDIA provide the lowest cost per token, uh, for example. Um, power utilization is very important. You know, essentially we're limited by power. Power limits um how much you can build, or what's going to be the size or capabilities of the factory, and therefore also how many tokens you can generate, and therefore also the cost per token. Um the network once we're once we understand that essentially the network is the element that determines what you actually build. Okay, the network and there is not just one network, there's multiple infrastructures, but the network determines if you have a server fund, that means that your power utilization is very low, and the cost of token is very high, versus building an Air Factory, versus building a supercomputer where you can generate many more tokens on the same compute, and the for the cost per token is much less. Okay, once you have the right network, it's dramatically impact the number of tokens you can generate per compute, which means it dramatically impacts the cost per token. Dramatically impact the the money that enterprises can make from the air factories in that sense. That's why the infrastructure is so important, and you cannot uh um just uh bring some infrastructure, some connectivity that is not fully designed with the compute, that is not fully synchronized, that it's not co-designed with the computer, extreme co-designed with the compute, because then you actually didn't build an air factory, you didn't build something that generates tokens that generate money, you build a data center that's actually will consume money in a sense, and do the and do the opposite thing of what you wanted to do. That's why infrastructure is so important, that's why networking is a key element that determines the cost per token. If you're at the right infrastructure, it continues to reduce the the cost per token even further and enable our customers and enterprises to actually uh generate more capabilities, more money, more tokens from the AR factories.
SPEAKER_00Very cool. So, Glad, we've talked a lot about how important networking is uh for these AI factories and so forth. And one of the things we're knowing is that we know now is that organizations are definitely getting out of the pilots and so forth. They're moving into scaling up their environments. And one of the things the research highlighted is that organizations were stating that they've really got some challenges with networking as they scale their GPU environments, right? Especially as the AI factories grow from, say, thousands to hundreds of thousands of GPUs. And a big part of that is also around resiliency, right? Because that's tied to uptime, which, as we were just talking about, revenue generation and the all-important return on AI investments. I'm wondering if you could talk about how networking is going to help maintain AI factory performance even during congestion, hardware failures, changing traffic conditions like we're seeing with agentic and so forth, without driving up any operational complexity.
SPEAKER_02Yeah. Um, and and you're definitely right. You have uh AF factories and they're growing in size, and you need to make sure that they're resilient and reliable and so forth. And uh, and and the full resiliency is an integral part of the entire AF factory design. It's an integral part of what we do on NVLink, which is uh very reliable infrastructures, and we build capabilities to continue and increase the resiliency of that. Um, the same on on the scale out. Um, and and if you look on um recent developments, for example, in infrastructures, one is moving to co-package optics. And co-package optics actually enabled two important things. One, it helped to reduce power consumption of the infrastructure. And as we know that power is a limiting factor, if you can reduce power consumption, you can bring more compute engines and actually generate more tokens on the same power envelope, okay, which is very important. Um, that's why, by the way, we're also um adding or added liquid cooling and 45 degrees of liquid cooling, so we don't need to have chillers and you save more power in other places. But on the infrastructure, co-package optics is a very important thing. Now, besides reducing power, uh co-package optics also increase resiliency of the factory. We're saying that with co-package optics we are increasing the resiliency by 10x. So there is an order of magnitude less interrupts versus uh pluggable options compared to co-package optics. So co-package optics doing two important things reducing power consumptions, so you have more tokens per an infrastructure per the power limit that you have, and we're increasing the the mean time between interrupts or increasing their resiliency by an order of magnitude with co-package optics. So that's one thing. The other thing, another example is that the the network topology is continued to be evolving, continue to evolve, continue to optimize. And if previously the the the scale out network was built in a um in a in a in a multi-levels of factory, for example, uh today we're building the scale-out infrastructure in a different architecture, in a new multi-plane architecture. It's an architecture that actually has a collections of factory topologies that are uh creating that scale-out network, which means that if you have a failure within that scale-out network, the impact of that on the scale-out performance is very negligible because you have a lot of other elements that works in parallel. So there are innovations in the topology, and this is sometimes that's people um uh kind of uh ignore that part, but the way that you connect everything, the way that you connect the switch to one another, the way that the switch is connected to the supernix and so forth, the topology of that infrastructure is very important. It's very important for the performance, it's very important for the resiliency. It's continued to be involved, and there are new innovations around that. And multipan is the new innovations in topology that's further increased resiliency and enabling to continue working and operation and running the operations of the AFACOR, even if there is a failure or temporary failure within the network itself. And of course, CPO, which is another um order of increasing the resiliency at an order of magnitude. So there is more uh innovations that are happening across the entire infrastructure, resiliency in around resiliency, making sure that the uh AF factories continue to work, even if there is a case of failure that uh will be fixed later on. You can still run and generate tokens and be very efficient.
SPEAKER_01You know, Gilad, we love having you on because we can do these deep dives and we could probably keep you here for another hour and get into adaptive routing and uh you know how customers should think about automation. Telemetry is a big deal, you know, the resilience of these networks. There's so much more we we could talk about, but but we really appreciate your time. I want to sort of ask you one more question, just sort of forward looking. When you think about you know, looking ahead on AI factories, you're gonna have massive, you know, training continues. There's this just distributed network that you described. You're now moving into real-time inference. You guys with STX have made some really interesting moves around storage. I think you've completely changed the way we think about that and the storage hierarchy, uh, agents and agentic and just billions of interactions. So, what does the next generation of Ethernet networking look like from your standpoint? Where do you see the big opportunities for innovation over the next you know year or two? Uh and even you know, further out.
SPEAKER_02Yeah, no, if if I'm gonna tell you everything now, then you know there's you're talking. And I'd like to meet you more. You know, there's more things that we can talk about. Um obviously, uh, there is a great amount of innovations that we have in mind and start to executing on and working with our partners and our ecosystem. I think the the the the main the major or the main uh new thing right now is uh the move to co-package optics. Um now NVIDIA uh Spectrum X uh co-package optics or Spectrum X photonics is um a moving to production and we will uh we have already a system running within NVIDIA. We've seen the great performance results, we've seeing the increased resiliency, the order of magnitude increase of resiliency. We see that we see the the power reductions. We're very happy with co-package optics that now move to uh full production. Um, we're gonna start seeing AI factories leveraging co-package optics and and so forth. So that's uh uh an important transition. Uh, happy to see that. It's a technology that was a couple of years in the making, and of course, with uh a great partner of us, TSMCs on packaging and and around laser resources and uh fiber arrays and so forth. There is a great amount of technology that uh we had to develop, we develop together with the ecosystem to make uh CPL a reality. Uh so that will be a major uh transition and happy to see that. Um, and of course, there is other things that we are looking at as continue to uh increase the capabilities of air factories, move to the next level of scale, and so forth. And I'll be happy to discuss uh those items with you in the next time.
SPEAKER_01Yeah, we'd be great uh to do that and love to have you back. Thank you so much. It's been a great discussion. Uh, I just want to say to the audience, you you know, we want to leave you with hopefully we've conveyed that networking is evolved from this sort of simply infrastructure that moves packets. It's now an essential part of the AI computing platform. It is a key element of the AI factory, whether it's enabling distributed inference, supporting JetTick AI. We talked about GPU utilization, um, economics, better performance per watt, which we really didn't get into, but actually Gilad talked about that. Lower token costs are gonna continue. Networking is increasingly central to AI success. Also clear that Ethernet has evolved considerably. And the debate about open versus closed, yeah, that's gonna continue, but open standards will evolve and be critical. We're seeing NVIDIA embracing them, intelligence software, tightly integrated system design, extreme co-design, all coming together, delivering scalable AI infrastructure. Galat, thank you for sharing your insights and helping us better understand where AI networking is headed. Thanks to you for watching. Check out more research and analysis on AI infrastructure, thecube.net, siliconangle.com, thecube research.com. I'm Dave Vellante, he's Pablo Libert Dal Liberte. Thanks for watching.