Semi Doped
The business and technology of semiconductors. Alpha for engineers and investors alike.
Semi Doped
WEKA's Val Bercovici: KV Cache, DeepSeek V4, HBF, SLC vs QLC NAND, CXL, NVLink, Tokenomics
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Vik welcomes Val Bercovici from Weka to discuss the rapidly evolving landscape of AI memory and storage. Val explains how Weka's architecture leverages high-bandwidth networks to make storage faster than motherboard DRAM. They dive into KV cache optimizations, the future of NAND flash tiers, and the role of CXL in AI inference. The episode concludes with a look at predictive memory offloading and the AI flywheel.
Chapters:
0:00 Welcome Val Bercovici, Weka
1:59 Memory situation and model routing
3:50 KV cache offloading to CMX
6:10 Network faster than motherboard
13:10 Weka as AI memory infrastructure
14:45 Inference market is different
16:06 Memory hierarchy and KV cache
19:40 KV cache optimizations and demand
25:20 DeepSeek's cache read pricing
34:49 NAND flash tiers: SLC vs QLC
43:01 High Bandwidth Flash (HBF)
49:59 CXL versus other interconnects
Follow Chipstrat:
Newsletter: https://www.chipstrat.com
X: https://x.com/chipstrat
Follow Vik:
Newsletter: https://www.viksnewsletter.com/
X: https://x.com/vikramskr
Follow Semi Doped:
Get more of Austin and Vik daily, free!
Sign up: https://daily.semidoped.com/
Welcome to another semi-dope podcast. I'm Vic from Vic's Newsletter, and with me is Val Berkovici, Chief AI officer from Weka. Weka is an AI data and memory infrastructure company. Val has been on Semidope before, and uh you know, a few months ago we spoke about context memory storage uh right after uh Nvidia announced their platform. So if you want to check out that whole conversation, uh it's it's right in the back catalog, and we'll try to link it up here as well. Uh but yeah, today's conversation is more about uh more recent developments in the world of memory and storage and where things are going in AI inference. Val, thanks for being on the podcast.
SPEAKER_01So much wanted to be to be back, Vic. And as we predicted, the only constant is change. So uh lots of updates to talk about since I was last on.
SPEAKER_00Awesome. Isn't that so exciting? We have like things that are constantly changing. And since we spoke, like, what was it, like a few months ago? Yeah. It seems like everything has changed. And you were mentioning before the show, seems like 10 years have gone by in AI land.
SPEAKER_01Exactly. In fact, I actually I don't know if I made this prediction last time. I expected a mythos fable moment at the end of the year, and it happened in like what, May or April, I forget exactly when. So things are definitely happening faster than we thought.
SPEAKER_00That's amazing. Yeah, yeah. It's it's really picking up now, and uh memory has reached an all-time high. It's become extremely expensive. I think a lot of people who do who are deploying AI hardware are like finding ways now to make the best of what they've got. Uh, increase utilization rates, pool stuff, offload to different tiers, uh, whether it's other kinds of RAM, not just HBM or go to NANDFlash. No, so there are so many techniques that have sprouted up in a matter of mere months because we we have to do something about the memory situation.
SPEAKER_01Exactly. I would say even one of the, I wouldn't say it's the simplest, but one of the most obvious and coarse-grained techniques in my mind has been model routing. If you're not always using the most memory-hungry model for every prompt, especially in an agent swarm with hundreds or thousands of turns, uh, and you're intelligently routing requests to medium and small models, implicitly there, you've got lower memory requirements. So that's kind of an easy button, if you will, for reducing memory in aggregate. But then we get obviously much more specific here with regards to each model's memory requirements, working memory, KV cache, and so forth.
SPEAKER_00Yeah, exactly. So you said that's what we were talking about in some sense last time, because the KV cache started getting so big when the agents arrived on the scene that uh it you couldn't keep it on expensive HBM because it's a very limited resource and it has to be preserved and used judiciously. So there was like a lot of talk about offloading to DRAM. Um, and then as agents started working more and more and having greater context length, they started talking about offloading to context memory storage, which is what was this whole rack, which Nvidia calls what was that, STX, right?
SPEAKER_01Uh yeah, the new storage platform is STX. The biggest use case of STX is gonna be context memory storage, CMX, context memory extensions in NVIDIA terminology. Uh, and that is yeah, within the STX hardware framework, hardware and software framework.
SPEAKER_00Awesome. Uh so yeah, I know Weka has uh hardware. We spoke about this last time too, where uh you have got uh Weka's AMG, uh which provides really uh a lot of storage directly to the GPU via a very fast network connection, and uh it occupies uh it the the speed that you get out of it is like uh really good comparable to uh uh something like uh faster than NAND but slower than DRAM, right? So that that's where we were.
SPEAKER_01Let's let's let's double click on that because it's actually uh different than that, even. So uh the standard positioning is you can be faster than NAND with higher bandwidth connections, you know, faster than local storage effectively. Uh but in a standard hierarchy, you could slot in a slower than DRAM. However, depending on your architecture and whether you can take advantage of the full line rate network performance or something like NVLink, the way Weka does, you can be faster than DRAM. So if you actually take a look at the number of you know PCI lanes that we have in these wonderful scale-up domains with NVLink on an NVIDIA servers in particular, there are far more PCI lanes, like 128, I think, in NVLink than there are PCI lanes from the CPU to DRAM, which is only 32 on the 32 on the motherboard. So there's more, this is one of the cool architectural, uh, I would guess, leverage points of Weka, is we recognized at company founding Inception about 13, 14 years ago, for high performance computing, which is what AI factories are, there is more bandwidth on the network than the motherboard. And PCI is a bottleneck, it's not an accelerant. And this is still something that I find a lot of systems architects and developers struggle with. They assume nothing is faster than what I can access on the motherboard, and they design entire database hierarchies, entire application, you know, latency profiles based on that invalid assumption today, because the network is faster than a motherboard in a GPU network and a high performance network. And so, yes, I think uh the memory tiering can have whatever you connect on the other side of that high, high bandwidth network, the NNV Link style network, you know, uh East-West, Compute, Memory Network, Nickel network, it goes by many names. Whatever you connect on the other side of that can be faster than DRAM. And the way Weka connects, uh optionally, we don't require that network, but we certainly leverage it maximally. Uh, but the way Weka connects, we're faster than DRAM on a raw bandwidth basis. Another confusing part is the way CPU and GPU memory transfers, KV transfers, for example, and others, operate on most of these servers, GPU servers. Uh depending on a generation, Hopper has something called a bounce buffer between the GPU and the CPU. Uh Blackwell introduced, especially with Grace Blackwell, more sophisticated chip-to-chip interconnects that are faster between GPU and CPU transfers. VRuben yet again continues to advance that with higher bandwidth, you know, far more scalable memory transfers between GPUs and CPUs, but you can still be faster than DRAM on these servers because of NVLink. There's a ratio in terms of NV link bandwidth relative to memory bandwidth relative to front-end network bandwidth. And as long as you can be line rate on NVLink, uh, you can be faster than DRAM, which is counterintuitive, but the math, math is math, it proves out.
SPEAKER_00That's awesome. Okay, that's great. Because uh just to recap, what uh what we are doing here is that the line the networking on modern hardware is so fast that when you hook up even storage using that, you can get uh bandwidths that are faster than DRAM because motherboards aren't the fastest. And that is an assumption that a lot of people are making right now. But instead, if you view the NV link as like a very high uh bandwidth interconnect, you can actually bring storage to operate with much higher capacity, but at speeds that are better than DRAM speeds.
SPEAKER_01Yeah. And uh these networks are really critical right now. We we've seen uh a rise in interest in networks, not just with Co-Package Optics, CPO, but what Google does with their Taurus network, AMD obviously with Pensando and that line of uh of networking that they're going to continue to improve with Helios. Uh, but just what Mellanox has done with NVLink has been wonderful. Uh and it's not just NVLink, right? Melanox offers CX7, CX8, CX9 network adapters, Bluefield 3s and 4s. There's multiple varieties of Bluefield 4s. Uh, and so we're not picky, at least at Weka, with regards to where the bandwidth comes from. As long as systems architects configure the bandwidth they want for the memory tiering they want, Weka can support it and uh and continue to deliver these radical economic benefits in terms of more tokens per second, uh more concurrent users per second, which is really hitting the market sweet spot right now of agent swarms.
SPEAKER_00Awesome. Yeah, so since we last spoke, like Weka right now shouldn't be viewed so much as like a storage company anymore. Because in the AI era, uh because of what solutions Weka provides, I think it's far more than that. We were speaking a little bit about it before the show. Do you want to double-click a little bit more into why Weka should be seen more as a memory company and a company that does AI data infrastructure instead?
SPEAKER_01Yeah, very much. You know, uh, two of sort of our industry partners, if you will, one is Grok with a Q that Nvidia Acqui hired last December for $20 billion and a non-exclusive license to their technology and most of their founders and engineers. And then, of course, a very uh very popular and successful Cerebrus IPO were clear indicators that uh the inference market, the inference infrastructure market, is very distinct, you know, very different than the training market that NVIDIA dominates with GPUs. And even Google themselves, for me, for some reason, this is like the clearest example. Google has had seven generations of TPUs that were applied both to training and inference, right? And radically different markets back then when you know TPU 1, 2, 3, etc. through 7 was released. With TPU 8, which they uh announced just a few months ago, they were very explicit as saying there is now a TPU 8 for training, but there's a completely different TPU 8i for inference. So that's just another example that you know uh inference is different, inference infrastructure is different. It's great when you can leverage the same infrastructure for training and inference. There's nothing wrong with that, but understand that you know, whether it's really cool companies coming back out of stuff like etched, you know, or or Cerebrus or all sorts of other LPU-based, uh, you know, ASIC-based, SRAM-based companies, inference will continue to diversify in terms of infrastructure and diversify away from training. So Weka is very much following in that trend. Weka has a great legacy of high performance computing, HPC storage. But the new tagline for the company, data and memory infrastructure, reflects the fact that at its first principles raw basis, inference is not storage-centric. Inference is a little bit compute-centric, as we know, for prefill and extremely memory-bound for decode and feed forward and so forth. And that's that's really where the Weka augmented memory grid product line fits in, is it's very much you can use it as an inference-only product without any Weka storage whatsoever. Uh, it just happens to be backed by NAND at the bandwidth of memory. And as we just discussed, when that uh when that particular context memory network happens to be large, then it's actually more bandwidth than DRAM.
SPEAKER_00Ah, yeah, that's awesome. So uh depending on the kind of hardware that is running now, you mentioned that Cerebrus is one, uh, and the whole inference landscape is like really, it seems like there's no one right way to do inference. You could run it all on SRAM, you could run it uh with like LPUs, which is also SRAM. Then you could uh run it with a combination of HBM um and SRAM, uh, or if you look at Samba Nova, they use all three. So it's like really breaking up all over the place. And uh so how do all these compare like in terms of performance, or where does storage play a role into this whole thing?
SPEAKER_01So there's a memory hierarchy, and it's interesting now because uh when we last talked about this memory hierarchy, it's funny. We we talk about it as if it's been around forever, but it's really only about a year old. But a fairly well-established four-tier hierarchy that's gonna have to change now. And NVIDIA's dynamo team has done a really good job of documenting this. They even label it G1, the top tier in the memory hierarchy is for high bandwidth memory. G2 is for CPU DRAM, but it also goes by uh LPDDR and SOCAM, but it's you know, for the uninitiated, it's all the same thing. Uh and then uh so that's G2. G3 is what's called local storage, and sometimes it's called like the RAC SSD that's in the servers. And then the G4 tier is remote storage. Often it's NFS or or S3 type storage layers. Uh, and the one thing to really not maybe confuse about this memory hierarchy is there's no smooth graduation from HBM to DRAM to storage, local or remote storage. These are very, very rough, sort of jagged transitions, uh almost a Grand Canyon sometimes between them, because we're talking about orders of magnitude higher latency between those memory tiers, orders of magnitude different bandwidth between those memory tiers. And you very quickly reach these cliffs where if you're okay providing one token or 10 tokens per second to a user and you're okay with latencies, time to first token of like 10 seconds and end-to-end latency of hours where they should be minutes, then you don't have a problem. But in the real world, no one tolerates, you know, five or 10 tokens per second output. Kind of the human eye demands human attention span demands about 35 to 50 tokens per second of output minimum. And agents, of course, at machine speed will take thousands of tokens per second of output in a multi-turn agent swarm. So you know, SLO's service level objectives matter. And that's where we quickly find right now that uh even though there's a lot of vendors talking about memory tiering, including NAND and storage in the KV cache hierarchy, it's kind of irrelevant. Uh, most of the benchmarks we see out there today that include storage in the KV cache tiering hierarchy are workloads that you could just run on a D on a DGX Spark or a Mac Studio today. You don't even need a big server to run some of these smaller models with small context windows and uh and you know single-turn chat sessions or two to three turn agent sessions that are not representative of reality. When you're running real-world workloads today, and uh semi-analysis is actually working on updates to inference X that they've nicknamed Agent X after my suggestion, actually. Um Agent X, when it comes out soon, I'm not gonna steal their thunder, but when that comes out soon, that's gonna reflect something called an input sequence length, an ISL, which is the context window, of much larger than 8K, which has been the limit in the past, hundreds of K. You know, they themselves have published uh analysis of the fact that I think the median traces uh for coding agents now are two, three hundred K tokens, because of course we're in the era since we last spoke of million you know token context windows instead of 100k or 200k token context window limits. So when you reflect the reality of large models, trillion parameter models, or models in that class, you know, Minimax is in that class, even though it's not quite a trillion parameters, but Kimmy is and GLM 5.2, the new hotnesses, and so forth. When you reflect large models with large multi-turned sessions with high concurrency, the way you know, Claude Code, Open Code, Codex, actual Pi, you know, Hermes, open cloud, actual agents work, then you end up with a radically different workload profile. Uh and you on the one hand, uh the models have gotten so much better in terms of KV cash consumption individually, right? Um I use a number I think I remember of for every 100k tokens with without optimized compression like TurboQuant, without optimized compression like sliding window attention and others from Deep Seek v4, you are up to 50 gigabytes of KV cash usage for that 100k tokens. That's been reduced on a unit basis right now with TurboQuant, with optimizations from Deep Seek and others, uh, to about five, five gig of instead of 50 gig. So that 90% reduction is great. But what's happened on the other side? Context windows 10xed, and then agents formed 10x, you know, or 100x'd the number of concurrent sessions, uh, and on and on and on. We've gone multimodal. Um, and so the ed the net effect we always seem to see is that for pretty much every 100x reduction in unit KV cash size, there's a 10,000 X increase in consumption of KV cash. Yeah. So we're always looking about that 100x more uh token volume, and it's reflected in the in the pricing pushback we're seeing by enterprises today. Maybe they're not paying 100x times more than they expected, but they're paying a lot more than they expected because of that token volume.
SPEAKER_00Yeah, yeah. So this is this is great, right? Because clearly KV cash optimizations are on on the way. Like TurboQuant was one which scared uh the daylights out of the market. Everybody thought like memory is dead, HBM is dead, everything is dead. But it turns out that the quantization uh is a good thing because we we only need, like you said, one tenth or 90% lower uh KV cache storage for 100k tokens. So 100k tokens was taking 50 gigabytes, now it takes five gigabytes. But this is great because which means with our existing models, we can run uh far more agents and do a lot more work. And so the demand really never went away. In fact, it got more, right? It increased.
SPEAKER_01Yeah, in fact, that that that it it's you could almost make a meme out of this. That 50 gig that went down to five gig went right back up to 50 gig because the context window size went from 100k to a million. So uh there's a give and take here, and the net result is just always more memory demand.
SPEAKER_00So just because KV cash is being compressed like this, do you see that it's going to continue to be so? Do you think that we'll go to like maybe one gig per 100k tokens in the future? And what does that mean for KV cash? Because if you say like it it we dropped uh uh usage by quantization techniques, and then we brought it back up because of like agents, I guess we don't have a net increase. Is that something that the memory and the storage industry should worry about? Because uh I don't know if the increase in the usage is exceeding the benefit coming in from the quantization.
SPEAKER_01That cliche this time it's different, applies here too, right? Every time there's this massive new innovation in KV cache compression, uh it's not different. It's just Jevon's paradox keeps kicking in over and over again, and we're seeing that oh, you know, that higher increase in demand every time there's a reduction in the unit you know consumption. And so for me, uh let's look at context windows. We're not stopping at 1 million. It won't be long before we're two and five and 10 million, maybe by the end of this year. Everybody wants more context. And uh we're going again, longer horizon in our agents. Uh, we're we're giving agents far more ambitious goals, being encouraged to by the agent and model providers. The goals now span hours and days. Soon, I think some agents will run weeks quite regularly. In fact, for cybersecurity, which we can get into, they run forever, right? You've got to run the Security Operations Center persistent blue agent swarms forever now. So we're we're seeing longer multi-turn horizons. Of course, we're seeing more parallelism, more concurrent subtasks. Claud workflows was one sort of feature introduction that became very popular to just solve a problem in parallel 10 times faster. And uh and then again, we're going multimodal. So it's not just language, but we're inserting more video and audio frames and image frames into these agents. Uh and it all just results in a continuous explosion of overall token demand and overall KV cash consumption.
SPEAKER_00Yeah, that's awesome. Okay, so the now that we are we have established that this is only going to get more from here, so NAND and DRAM or whatever like form you would want to look at it as pooled, not pooled, we're gonna need more of this because there is, we are only going to run more of these things, so we're gonna need more. So that doesn't spell any decline or leveling off uh in the near term because context windows are gonna get longer, agents are going to run longer, and more agents are going to run per task because you can like have parallel ones run. But I think one interesting thing that you mentioned earlier was like Deep Seek's optimizations, because uh when Deep Seek V4 came out, there was a lot of emphasis on uh SSD first approach to inferencing, which I think um really helped in terms of their token pricing, especially when you end up hitting cash. And they showed that for some workloads you could hit like 95% cash hits. And what happens? Then is that as long as you keep hitting KV cash, uh you already have the tokens stored in a high bandwidth network connection to an SSD? Which means now you can uh basically get like free. It's free now because they their token pricing is so low for cash hits, uh, it's it's insane. Like, how is this playing out into like other models? Or are other models also hitting cash hits at this level? Or how what do you think of that?
SPEAKER_01Yeah, um so on the one hand, I'm glad you noticed that because we really want to emphasize this key point now. Uh, the majority of token volume is agent swarms by far, 80-90%. It's no longer chat sessions. So chat sessions are almost irrelevant to the conversation from an infrastructure perspective. And um, and agents, uh, it's interesting, even uh the Sonnet 5 announcement just the other day, right? That everyone looked, even Anthropic, interestingly enough, said here's the new pricing for input tokens, and here's the new pricing for output tokens, and it's a promotional price. And the reason why I don't even quote the number is it's really irrelevant because of all the tokens used by agents, most of them are cash read tokens. And some of them with anthropic's particular, you know, um unique uh pricing models. You have to pay for cash rights as well to benefit from cash reads. But we should only be talking about cash read pricing. The other pricing is a rounding error, and Deep Seek's pricing reflects that, right? They're dramatically new, you know, X inserting two zeros, right? Basically after the decimal point uh for price per million cash read tokens really is a shock to the industry. Not so much the memory industry, it's a shock to the other models and the other inference providers. But the very important asterisk there is that pricing is only available out of China, right? You have to use Deep Seek hosted, you know, the rumor is, I think, a Mongolian data center with very, very low energy costs, but also some secret sauce that that DeepSeq has with regards to how they've integrated version four Flash and Pro with the uh you know uh fat high flyer file system, the 3FS file system, parts of which is open source. They famously open sourced that last year and wrote some wonderful papers and blogs about it. Curiously, unless I've missed something, they haven't updated those blogs with what definitely are some new innovations and improvements that factor into that dramatically low pricing. The net of it though, uh I think VentureBead did some math, uh, which I was quoted on earlier on, is it's about 87 times lower cash read pricing from China than the same model, via V4 Pro or V4 Flash, hosted in Singapore or hosted in the US or Europe. And so the real conversation is yes, DeepSeek has completely set the bar for global pricing for people that are able to inference out of China. But for people that can't inference out of China, uh, it's still by far the best pricing. But there's a massive margin opportunity for open weights inference providers to differentiate on cash read pricing because you you may not be able to match Deep Seek's subsidized pricing, but you can still leverage their innovations, their sliding window attention, hybrid compressed attention, HCA, and so forth. You can still leverage those innovations and offer very important uh reductions in token cash read pricing and really drive effective, you know, aggregate agent pricing down because that that is the one pricing metric that dominates the cost of agents.
SPEAKER_00Yeah. Uh so I have like a couple of questions on this one. The first one is why is it I it's news to me that that pricing comes only out of China. So if I run a deep seat model out of, I don't know, open router, I don't have that pricing.
SPEAKER_01Open router is wonderful because it shows you the difference, right? In fact, for some reason, I don't even know exactly what this signifies, but there's a little um a little slider button you have to slide to say show ignored. And I'm not exactly sure why they label it button the button that way, but when you click on that, it does expose the deep seek pricing out of China. And then you can compare. In fact, your agent, of course, your Hermes, your OpenCly agent can compare uh side by side in real time the pricing for the exact same model, the exact same input pricing, the exact same output pricing, the exact same cash read pricing. And I should double check because I I tend to stand that open router site quite a bit daily. I should check whether there's more and more cash write pricing beginning to appear, because I know for some models it has, and uh it escapes me whether it's appeared for DeepSeek providers as well. But you're gonna see some providers now differentiate, not just on cash read pricing, which is a major point of differentiation, but also on cash write pricing.
SPEAKER_00Okay. Uh so on term in terms of like cash read, in what the real workloads that we're running today, it could be like agent work agent workloads, right? Because, like you were saying, any other kind of workloads like chatting and typing in, you know, uh on the chat window is not a negligible portion of the inference market today. So let's just talk agentic workloads. Uh how does somebody uh make sure that they have a cash hit rate of uh I don't know, 95%? Because that will drive the pricing down enormously uh and create meaningful differentiation, like you were saying. And where is the industry right now in terms of cash hit rates?
SPEAKER_01Yeah, you know what, we could spend a whole pod just on this question because it's um it's very opaque and a little bit confusing. So let me uh let me go through this because uh this comes up a lot in other pods as well. When you look at your um agent dashboard, your clawed code, or just your clawed dashboard, if you have both code and co-work and other clawed products, or open claw, pick your dashboard, right? It'll have a cash hit rate. That is a logical cash hit rate. That that reflects the cashability of the tokens in your agent swarms. And that is often very, very high. Like every agent pretty much has a cashability of 95%, right? Because you are reusing a lot of contacts for agents, especially agent swarms and so forth. However, the providers, you know, the the actual cash hit rate from the provider uh is not one-to-one the cash ability of your tokens, right? The effective cash hit rate is very much a function of the memory tiers you have. And everyone, or by and large, you know, uh most people in the first quarter of this year and some in the second quarter have had very finite fixed memory tiers. There's only so much HBM that comes packaged on your GPUs, there's only so much of the RAM that's on the GPU servers. And as I mentioned before, you can try and add storage tiers and KVR floating to storage, but they ruin your SLOs. So it's really not that quick common in production for the popular models and popular token consumption. Uh, and so the way to really understand where where people actually have true effective cash rate rates that are as high, maybe not as high as the logical cash rate rates, but close is in the pricing. So open router, again, is a great proxy for that because it's the only way to introduce transparency to the real-world cash hit rates. And even open router themselves, they publish actual by provider, by model by provider. They publish some cash hit rates. And I like them because if you look every few minutes, they do change, right? Based on the actual token traffic of the moment. Uh, but even with some of the numbers I see there, I don't think they're the real effective infrastructure, hardware, memory, cash hit rates. There's some blend, but nevertheless, they're much closer than what your own your own agent dashboard reflects. And uh, and yes, uh for me, I like to joke now. We're we're seeing a lot of you know benchmaxing in a lot of the models right now. You know, for example, like Sui Bench, everyone sort of trains for, so it's no longer that relevant to benchmark, but deep suite is still a good benchmark. Uh, we see the same thing, obviously, when when when vendors sort of you know benchmark their KV offloading solutions, there's not a lot of truth left in benchmarking. So, my personal belief is that profit and loss, right? Pricing, transparent pricing is the ultimate benchmark. And uh, and you're gonna see Weka be more explicit in that space to do that, right? We we want to continue to prove out our own advantages, and we think that you know, real-world metrics reflected by pricing is probably the best way to actually benchmark a solution now, as opposed to something in a lab.
SPEAKER_00Uh you also mentioned that uh deep seek's uh sliding window attention is something really to look into. Uh what's what's unique about that uh sliding window attention in DeepSeq?
SPEAKER_01It's actually uh if you've ever paid attention to where compression of files in general has evolved from simple like PK zip compression or simple block-based deduplication towards these similarity hashing algorithms and a combination of you know, sort of global similarity and global hashing and local hashing. The same thing is happening now with KV cache. The same concepts are being introduced at the token at the attention level. And sliding window is just that. It's a way to take a look at just the more recent tokens and compress them as much as possible and just assume the tokens that your particular attention head hasn't paid much attention to recently aren't as relevant, aren't as compressible. And it's it's just cumulative, right? I think if you count it, there's uh formally about five different attention mechanisms in Deep Seek V4 now, all simultaneously applied. And each one specializes in you know short-term compression, long-term compression, you know, uh context-relevant compression and so forth. And the net result is that impressive 90% real-world reduction in KV cash usage per request. But again, the requests now are just coming faster and bigger. So it's kind of uh necessary to support. In fact, I think the reason why they continue to innovate so aggressively on KV cash consumption is they want to support million token context windows in soon two and five and ten. And the only way to do that effectively is to keep being more and more efficient, intelligent, current, if you will, on how you consume KV cash.
SPEAKER_00Okay, yeah, yeah. So uh Deep Seek, what it does, the sliding window attention idea is that you just keep the window of attention to what is most relevant right now and keep discarding what stuff is isn't relevant anymore. And that saves you KV cash.
SPEAKER_01Almost assume what's what's been attended to before that's not being attended to now has already been compressed to some extent. So now let's really be aggressive on what we're tending to right now and see where the opportunities to compress token memory, you know, is token attention in KV cash is.
SPEAKER_00Okay. In terms of just like NAND flash storage, even I I wrote an article recently about uh what it really takes uh for an SSD to be like AI ready. Like SSDs in the past uh were meant for a different uh use case really than what they're being used for now. And one theory I have, which I want to run by you, is that NAND itself is now breaking up into several tiers. Like we always in the past wanted to get more capacity out of NAND. So the industry was like trying to go from single-level cells, which has lower capacity, but you know, more endurance and um uh faster write speeds, uh, to going to QLC or you know, quad level cells on the other hand, where you can save, you know, so like basically store four bits or 16 different states in just like one cell. So it really gives you a lot more capacity, but endurance is a concern, and uh, I'd say it's harder to write to because you have to make sure you can differentiate between 16 states. How do you see these NAND tiers working out? Are single-level cells making a comeback? Uh, how is that working?
SPEAKER_01Yeah, it it is making a comeback, and so uh just like we talked about earlier, uh you know, training infrastructure used capacity NAND because you wrote a lot of training data infrequently and you read it very, very frequently. So that was ideal for QLC type approaches. Whereas inference is just fundamentally different, right? It's really fundamentally memory-centric. So you're really trying to make NAND appear more like memory and a lot less like storage. And memory doesn't care whether you write or read at the same rate, because the assumption is there's no endurance issue, right? There's a power issue, but there's no endurance issue, a persistence issue, but no endurance issue with with with memory, uh, with with memory, with DRAM in particular or HBM. Uh, and that's not true. So this is where NANFlash struggles to act more like true memory from a performance and a consumption perspective, as opposed to just a capacity perspective. Again, something this is something that Weka anticipated a long time ago. Uh, and so you're seeing that largely because of inference alone. Uh Intel had this really interesting technology with Micron called Optane or 3D CrossPoint. It's a real shame. It's a real shame.
SPEAKER_00Oh my god, yeah.
SPEAKER_01You know, today inference is the killer app for uh for optane, but it's gone now, right? Based on different materials, phase change memory, and so forth. Uh and what is try to replace it is SLC flash, right? Okay. If you have to, if you can't optimize your rights, your KV cache rights, uh well enough, then you have to buffer that endurance issue, that uh that driveware issue, with a more uh endurant form of NAND flash, which is a single layer cell SLC tier. And sometimes that can be complemented by QLC. So you buffer a lot of rights and SLC and then you destage them down to QLC later. Again, that adds complexity, that adds latency, that's no free lunch there. Uh another thing you can do, and this is something that Weka does, is kind of anticipate that you're gonna write to NANFlash a lot, you're gonna read to it a lot, basically use NANFlash as opposed to just keep it for capacity. Yeah. And in that case, yes, um, what Weka's done is we've always amortized rights across a whole fabric of NVMe drives. My joke is there's no S, there's no storage in NVMe. It's non-volatile memory express or memory extensions. And when you treat it like a true memory protocol and you amortize the rights across a whole fabric of NVMe devices, you don't need SLC. Uh we don't have SLC, for example, never have in Weka. We literally can use TLC for a very high write endurance workload like KV cache. And uh of course, in in the future, as we're able to uh influence the uh the rest of the inference stack, so the scheduling and the routing, not just the KV transfer storage layer, uh, we will be adding support for QLC drives as well. Uh but again, we you gotta be careful about that because you don't want to break those drives by doing unintelligent things on scheduling writes or unintelligent things on just routing tokens to somewhere that needs a quick write, but is really kind of not worth it.
SPEAKER_00So you can actually optimize uh the controller layer and make sure that you don't clobber the drive, whether it's a TLC or QLC drive, and intelligently write to the whole TLC slash QLC array and still keep endurance.
SPEAKER_01Exactly. This is uh something that's a category of storage technology called shared everything, as opposed to the Hadoop style shared nothing, which you still see in parallel file systems like GPFS or Luster, et cetera. And with a shared everything approach, the client, if you will, has global awareness using a RAF protocol of the Q status, the work queue status. There's multiple queues in every NVMe device, every SSD controller basically. So when every client has global visibility into the queue depth of every NVMe device in a fabric, tens of thousands, hundreds of hundreds of thousands of cues on tens of thousands of drives in a fabric, then you can treat NVMe like a cache line for DRAM. And you can be very intelligent about where you you um you know you you shard the rights, right? Where you load balance the rights very, very intelligently. And at that point, you don't need expensive SLC tiers to buffer rights. You're you're not being you know suboptimal, you're not being unintelligent or treating devices as blunt storage devices. You're really you're digging into the devices and you're making sure that you're you're you're optimizing the actual underlying NAND flash well and cooperating with the controllers versus offloading work to the controllers.
SPEAKER_00So you don't really see SLC as coming in on its own tier. Maybe it's just going to be used as a as a buffer.
SPEAKER_01Well, let me be, you know, let me disclose here that I'm talking about how Weka has decided to optimize flash tiers. I think the industry needs SLC, right? SLC. The industry doesn't have our patents, doesn't have you know our implementation by and large, uh, except for the server partners that we partner with. But um generally, uh that that technology is not available outside of Weka. So you do need SLC. There will be a rise. You know, NVIDIA is forecasting this themselves, even for CMX. There will be a rise in the need for SLC to buffer the rights when you can't amortize them over TLC. Uh, but the the upside of that is that if you buffer correctly to SLC, you can still use QLC. At Weka, we're skeptical, right? Because uh there again, these are not capacity workloads where you have the luxury of time to destage SLC to QLC. These are very bursty, very intensive workloads that are attempting to emulate memory. Uh, and so we think that ultimately one tier, whether it's SLC or TLC, is the only safe and fast way to offload KV offload to storage. But you know, the industry, there's a lot of smart engineers across the industry and many vendors. You know, the market will prove what works over time. We just know what works for us right now.
SPEAKER_00So in a NAND storage uh uh rack, uh, you know, you could amortize rights over many drives, but do you also overprovision uh like uh and with to what ratio? Like if you have a hundred drives in a in a rack, and to make sure that you have the right endurance, even after amortizing over TLC drives, do you actually have 20% more? What is stuff?
SPEAKER_01These these are tricks we can always play, right? So if we don't have SLC and economically or just supply chain issues, we need to use TLC or God forbid today, QLC only, then yes, you would really have to overprovision. And the the net effect is you you would be buying, you know, let's say a petabyte of storage, but only use only see 300 terabytes of like effective or 500 terabytes of effective capacity from that petabyte that you've purchased and are powering because the overprovisioning is needed not to break those drives.
SPEAKER_00Yeah, so it could be even a two or three to one over provisioning.
SPEAKER_01Yeah, it's it's it's tough, you know, because these are extreme workloads. There's there's nothing gentle about emulating memory with NVMe with with NAND Flash, right? It's a very intense workload.
SPEAKER_00Yeah, yeah. Well, that's great insight. That's uh really great insight for me uh on how SLC and you know TLC tiers work and what the trade-offs are. Uh so I I definitely has added to my uh body of knowledge today.
SPEAKER_01So it's nuanced. That's why I said you know, we could spend a lot of time on this alone.
SPEAKER_00So yeah, yeah, yeah. Well, we should move on. I think like a lot of interest in the market right now is for high bandwidth flash, because now if people can put in uh you know uh flash right next to uh a GPU and have it store some stuff, um, it would be a useful thing, I would imagine. So, my two questions I think around that are does HBF use SLC, like single layer cells? That's one thing. And secondly, what are really the use cases for this thing?
SPEAKER_01Yeah, great question. I think uh HBF will probably have to use SLC in the most common configuration for the reasons we just discussed. It really is trying to be a great offload tier. Uh it's trying, in fact, in many cases, to replace DRAM, right? Right. And complement high bandwidth memory directly. Uh for me, it's always come down to a packaging decision. You know, I think SK Heinex has been public about the fact that they want to package it, you know, directly on the GPU package itself. They don't want to go through buses, they don't want to go through networks, they want it to be really tightly bound to GPUs. I think it's a great vision. I haven't seen any GPU vendor uh directly commit to that yet or adopt that yet, but I think uh in the future, maybe 2028, I think we'll see that. Uh, but there's other vendors that I can't discuss, Andre NDA, that are packaging it differently uh on a PCI card, for example, and not even packaging it with GPUs, but with ASICs, you know, non GPU accelerators. So there's gonna be different packaging form factors, but ultimately I think it's a necessary thing because the opportunity. To get this right, the opportunity to optimize the KV transfer layer, the opportunity to schedule correctly, it was really acquisition, I think, a Qualcomm of Modular, right? And uh basically, yeah, and be able to get that compiler expertise. Uh, what NVIDIA acquired with Grok was a lot of compiler expertise from the TPU team over at Google. That is going to be more important over time to making HBF really usable because as we recompile models to understand HBM and HBF tiers, and maybe bypass, you know, this is a prediction we have inside Weka engineering, bypass DRAM tiers altogether because you can get a lot of bandwidth out of NVMe flash, you know, uh NVMe uh sorry, out of NAND flash devices over NVMe or other protocols or without NVMe, but just NAND flash, uh, there are optimizations possible now, and there's definitely a trillion dollars of opportunity that encourages those optimizations between HBM and HBF with RAW, NAND Flash and or NVMe.
SPEAKER_00You can always replace drives if you can run that bandwidth over a high-speed network and NVMe drives, and like we spoke about over provisioning and something dies, you change it out or whatever. HBF, I I don't know, the one thing that always is on my mind is it is NAND flash after all, and it will, even if it's SLC, it ultimately has an endurance uh to it. You can't change it out if it's packaged next to the ASIC or GPU. I agree. So is that like something that's going to play out, or the lifetime of this thing is, I don't know, 50 years, doesn't matter. We probably could throw the chip away by then anyway.
SPEAKER_01You've already done a good job in covering this. There's a bag of tricks engineers are throwing to mitigate this problem. It's not one solution or one workaround. There's over provisioning is one of the tricks for sure. Uh then again, there's proper amortization just at the storage at the NVMe layer, NVMe fabric layer. Then there's very explicit scheduling at the inference at inference time, but also very explicit scheduling at recom you know, model recompilation time and very intelligent token routing inside the MOE level and even just at the gross model level as well.
SPEAKER_00Okay, okay. So you could engineer the whole thing to a level that this is not really a problem, uh, and you can still use it.
SPEAKER_01Not only can you do that, you know, uh I've already seen anecdotes of people applying Fable, right, to very complex engineering problems and seeing, you know, months worth of sophisticated systems engineering completed in four hours. So one optimistic scenario is that these agent harnesses and the eval loops, the quality eval loops and so forth, the judge judgment models, and of course the guardrail loops and the guardrail tokens, when you package that all together, some of these really deep tech engineering advances can happen much faster. But ironically, we need more engineering in the harness to make sure they happen reliably and safely.
SPEAKER_00Okay, that's amazing. Yeah, it's so much stuff is happening that like we it's really interesting to see how all this plays out. So that's that's the whole fun of this thing. And the one other thing that has recently cropped up is uh, and I know like we like we spoke about this last time, and uh um you had like ideas around this that I want to see if it has evolved or changed ever since is the use of like CXL. There's a lot of talk about using CXL now because some of the RAM, DRAM isn't really being used on every server. I think maybe like 50% is being used. So because memory is so expensive, uh even Google has had a change of heart, it looks like, to actually start using CXL and reclaim some of the unused DRAM. Uh what do you think? Like, is CXL a thing now?
SPEAKER_01Or so CXL has a lot of fans, and I was one of them 10 years ago. Uh I'm not a fan anymore, right? And it's not because I don't like the technology, it's because in the real world there's alternatives. And so um, you know, personally, I think Mellanox and of course Nvidia's acquisition and really great execution of scale-up domains kind of killed CXL in one sense. The ability to have this great NVLink style network and have it used for memory, have it used, you know, um, for high bandwidth as well as regular DRAM memory uh has been, you know, a real, you know, uh glass ceiling or or concrete ceiling, really, for CXL growth, market growth. There's been less and less need. Then you've got Rocky, right? RDMA over Ethernet, uh, and you never bet against Ethernet in this industry, right? So the challenge CXL has, it's another bus. It's another bus you have to engineer, it's another bus you have to debug, it's another bus you have to maintain and power. And it's not that in isolation it's bad. And again, it's got a great killer app today of utilizing, you know, this really precious underutilized resource in some cases of DRAM. But uh in the context of real-world alternatives, I'm just pessimistic about the future of CXL. Uh, largely because, again, I'm I'm biased. I'm able to leverage things like Rocky or NVLink over Infiniband and deliver better than CXL performance with NAND flash economics, cost of goods, and capacities. So it's one thing to pool underutilized terabytes of DRAM, it's another thing to pool petabytes and exabytes of NAND flash at the same or better performance, right?
SPEAKER_00Yeah, won't uh won't a faster network bandwidth also benefit CXL like it does in LAND?
SPEAKER_01The first principles are simple until you have to engineer the real world issues into it. But uh yeah, the first principles are you know, more bandwidth is good. And so if you can have uh faster than PCI bandwidth to a pool of DRAM, as Google discovered, there's benefits to that. It's just that uh if if there really were benefits to that, I think it's you'd have seen it in Blackwell, you'd have seen it in Vera Rubin from NVIDIA or in Helios from AMD, or even in Feynman, you know, the the the next generation from NVIDIA that's already pre-announced, and you haven't seen it, right? And then you haven't seen it as a standard supermicro offering, and you haven't seen it as a standard Dell or HPE or Lenovo offering. And so I think it's that absence that speaks volumes, right? Uh there definitely are use cases for it, but for some reason it hasn't broken out into mainstream.
SPEAKER_00Yeah, yeah, for yet, probably. Uh it seems like uh yeah, it seems like uh the the shortage of HBM is now causing some players like Google to actually like consider it now because it's just like capacity issues are pushing uh you know hardware makers towards solutions they probably didn't consider before. So yeah, that's another interesting thing to see how it'll play out. Because you're right. Like so all this time it hasn't been in there. There's a reason for that, right?
SPEAKER_01There's a reason for that. And I think there's one important hint, right? So it was reported, I think, by semi-analysis about two or three weeks ago, that NVIDIA changed the bomb on Vera Rubin. And they cut the actual amount of DRAM in half, either from two, four to two or three to one and a half, depending on the model. But that's a clear indicator that A, DRAM has gotten too expensive, and B, uh NVIDIA is projecting with CMX solutions in the marketplace from a number of vendors, including Weka, that there will be less need for DRAM for inference. And uh, and so you're seeing that you know people are applying solutions to this problem, and it's not always just pooling it better, it's just reducing it overall.
SPEAKER_00Yeah, yeah, yeah. That's a part of what we spoke about in the beginning, right? Like you could make optimizations to the algorithm, you could use sliding window, you could do a whole lot of different things uh that uh, you know, maybe uses less memory overall going forward, rather than just use existing architectures but start pooling stuff together. So yeah, you you know it could go either way. If we find better algorithms, then we probably don't need to pool it. Could be.
SPEAKER_01Yeah, yeah. That's that's I think the prediction we're making.
SPEAKER_00Yeah, speaking of better uh algorithms, what do you uh what's your take on AMD's uh Mext acquisition? Uh and for people who may not have heard of this, uh Mext is essentially a software company that uh find found a way to optimize the uh use of DRAM uh by uh dynamically offloading all the unused parts of DRAM to NAND flash. And then uh using AI to predict when that same information is going to be needed back in the DRAM and preemptively moving it back before the GPU even notices it's gone. Right? So it looks like it's it's a cool way to use NAND, but what what's what's your uh engineering interpretation of what's going on?
SPEAKER_01That was a fun one to review because I wasn't familiar with MEX beforehand, but it was clear in reviewing uh their use cases pre-acquisition and the initial positioning post-acquisition is it it's it's AI technology, it's machine learning, specifically, small language models and small neural networks, uh in optimizing cache, you know, caching algorithms and being more semantically aware than just lease recently used or so forth simplistic heuristics. So it's it's a the application of machine learning and deep learning into caching algorithms, but the actual use case is ironically enough, not yet for KV cache offloading. It has potential to be very good there. We've seen even uh you know things like um you know popular in-memory databases, Redis and so forth, uh implement algorithms that uh you know aggressively destage, dramp to NANDFlash and so forth and retrieve it back. And Redis does position that for KV cache offloading as well. So I think it's an active space. Uh but right now, I think you know the AMD initially is targeting scientific computing. So whether it's life sciences, whether it's Monte Carlo simulations and other kinds of you know, seismic analysis, weather analysis, those are the kinds of applications that that technology has proven itself in. And it we may see it happen, you know, we may see it appear in KV cache offloading as well.
SPEAKER_00Nice. Yeah, I I thought it was an interesting use of the predictive nature of an LLM because uh if you can predict what the next word is, uh why why not use it to predict what the next uh page of memory is required and quickly pull it from flash to DRAM? Uh I I don't know how real uh practical or useful it is, but I found the idea was interesting.
SPEAKER_01Yeah, I think your instinct is right. You know, why why has cursor customized you know KMK 2.5 into Composer? Is many companies are realizing now that the the uh the the bar towards being able to train your own model has come down. It's a much more accessible thing for many companies right now. You don't need million-dollar ML researchers to train your own models anymore. And when you can customize a model for your domain, it actually doesn't have to be an LLM at all. It can be an ML, it can just be a very tight neural network. It can be inferenced on a CPU, it can be inferenced on a small, low-power CPU if the model is really domain-specific. And uh it's just a neural network at that point, it's not a large language model. And yeah, there's, I think, going to be again another Cambrian explosion of use cases and applications for clever small models that do things that heuristics, you know, uh peaked at and can no longer optimize or improve.
unknownYeah.
SPEAKER_00Do you think that this uh use case is basically physical AI and robotics, uh, where you could have those sensors uh at the edge like process very specific amounts of information? It's only one kind of information from a sensor, right? So it's not like a large model you need. So is that is that a useful use case you think going forward?
SPEAKER_01100%. I think it's probably going to be the reference architecture for robotics and edge inference. Uh we don't need large language models. Uh, there will be some aggressive, you know, uh cloud connection, whether it's through Starlink in remote locations or just broadband, if you really have to burst to some kind of complex decision that a large language model has to make. But you know, 90, 95, maybe 99% of inference for robotics will be local and disconnected, air gapped, so to speak.
SPEAKER_00Yeah, that's uh you the fact that like you could make uh inference decisions at the point of sensing, uh, or you know, just like put an intelligence anywhere actually is a very very useful edge use case. How useful or how good that intelligence is, I think, yet to be seen. But in principle, you could deploy these little models like everywhere.
SPEAKER_01Yeah, the cost of these Raspberry Pi style system on a chip motherboards are are really plunging. And uh fortunately, again, whether it's KB cache optimizations or just you know small model quantization and and just custom neural network training that doesn't have to be a large language model at all, is intersecting really well with really affordable, you know, uh system on chips, so cs. And uh yes, that results in some really interesting robotics and drone use cases, yeah.
SPEAKER_00Yeah, didn't Jensen mention something about the AI flywheel on in this context?
SPEAKER_01Exactly. So this is a general concept of uh it's all about being able to capture data, domain-specific data, train models, and then customize that with more either real-world domain-specific data or synthetic data now that you have enough real data to create useful synthetic data, and just keep iterating on that loop of you're pushing the frontier with big models, you're customizing either through fine-tuning, distilling, you know, quantizing, low-rank adapting, et cetera, all sorts of customizations, you're customizing smaller and smaller versions of those models. You're able to maybe retrain entire small neural networks that are very domain-specific of those models, uh, and and just get more and more efficient at processing inference, retaining some of the new fresh data, and keeping the flywheel going. So it's a it's a mix. It's definitely a whole ecosystem. It's a thriving ecosystem of different model types, different phases of data, different types of data. But uh, if you keep the flywheel going, it stays relevant with uh nature's natural entropy. So, yes.
SPEAKER_00Yeah, that's it's a fascinating idea. Uh, since we're coming up on time, I want to pick your brain for like one prediction. What do you think is going to happen in the next 12 months? What are you most excited about?
SPEAKER_01So, zooming out a bit, uh, again, this is something Jansen referenced very, very often, is that software is fundamentally changing. You know, um, a year ago, we wouldn't have predicted that all of our engineers really would be using AI for most of their daily work right now. It was it was heresy even a year ago. Uh, and so the not only the rate of change that's happening right now, but the fundamental change in software is that more and more software now is not compiled and run. It's basically compiled, run, and inferenced, right? It's really agents now, as we said, being much more intelligent in their token consumption. We're not using Opus for everything, we're not using GPT-55HI for everything. We're definitely using uh now model routing in a mixture of models and a mixture of experts within models right now to be very token efficient. And what that means is now the cost of running a software business is radically different than before. It is a high marginal cost. You can't just leverage, even with, you know, maybe KV cash is the way to leverage, but you can't just leverage the cost of tokens across users the way you could leverage, you know, cloud instances and databases and VMs and micro VMs and containers across users. Uh and SaaS companies, the reason I believe the SaaS pocalypse is real and is a problem, is that no matter how much SaaS companies figure out, you know, new pricing models and new values, you know, value-based pricing and so forth, their opex is going through the roof. Their opex now is token opex. It's tokenomics, it's token consumption. And yes, all these engineering solutions we just discussed are ways to manage those. But if you just take a look at token volumes on OpenRouter, week after week after week, it keeps rising and rising. And we've really just barely begun mainstream token consumption and persistent agent swarms. Uh, the only way to run a profitable gross margin business in software will be to own more of the token stack. And you can continue to outsource that to an inference provider, to a model provider, or you can acquire it. You can merge with a neo cloud, you can merge with a token factory, and you can essentially can you know vertically integrate more of that very expensive token generation stack uh and continue to run a high gross margin software business. Oh, that's fascinating. The real question is with cash flows, what they are, will SaaS giants acquire Neo Clouds before Neo Clouds are able to acquire the SaaS giants?
SPEAKER_00That's fascinating. You know, I always spoke about like uh the best way uh for large companies to save on token costs is like you bring inference on premises, right? Exactly. Uh so you could run that at the edge. What you're suggesting is like one level, that concept on steroids. A big enough software company can go acquire a neo cloud and say that this is my token factory, and now I can like there's still a cost to run the token factory, but the tokens are yours to use. It's entirely yours. Now, if every software, big software company starts acquiring neo clouds of some size, I mean, they don't have to be like multi-gigawatt data centers.
SPEAKER_01I was predicting, you know, I made this prediction before X acquired Cursor or SpaceX AI acquired cursor. I predicted, you know, like Workday or Monday.com or something like that would merge with like a mid-tier neo cloud. But now, of course, first dominoes fallen with X Space X AI and Cursor, and older established SaaS companies are going to have to react as well. So, yes, I think whether the dominoes start falling in the middle or one side or another, it's kind of inevitable now that most SaaS companies and most NeoClouds will have to merge.
SPEAKER_00That's a fascinating prediction. I would love to see how that works out. Yeah. Yeah. Uh thanks so much, Val. Like it's always a pleasure chatting with you. You're like a fire hose of information that I know I'm going to listen to this podcast later myself as I'm reviewing the edits and stuff and be like, oh my God, I missed that when I spoke to Val. But you know, I hope this helps all our viewers as well. It's really a pleasure.
SPEAKER_01Always a pleasure. We said we'd enjoy it the next time, last time we did it, and I'm definitely looking forward a few months from now from coming back. We should come back, but well before the end of the year, because by the end of the year, again, we're going to be very surprised by what happens.
SPEAKER_00It's an eternity. Every three months is an eternity in AI time. So yeah, we should do this more often. All right, guys. Uh, that's it for today. Thanks for listening. Uh, if you're enjoying semi-doped, please share it with your friends. And we also have a daily newsletter on semidoped.com where uh we put our daily takes on the news. It helps us keep abreast of what is happening in this fast-paced AI uh landscape we are in. And it's entirely free, so make sure to check it out. And thanks for everyone who puts comments on YouTube. We do read all of them. Some of them are like really amazing, some of them are really funny. We have a good laugh. But we read all the comments even if we don't respond, we promise. And it helps us plan all the future episodes. So definitely keep them coming. And if you can, leave us a five star review on the Apple Podcast, it really helps us out. Right. Cheers and catch you on the next one.