Speaker 5

Hello and welcome to a weekend news episode of the Leveraging AI podcast. A podcast that shares practical ethical ways to leverage AI to improve efficiency, grow your business, and advance your career. This is Isar Matis, your host, and we have a lot to talk about this week First of all, we have a AI agent that went rogue and hacked Hugging Face's databases. This is a really scary development, uh, but we have a lot to cover on that. We are going to talk about the age of routers that is now at full swing. There are routers popping up everywhere, and more and more companies are investing more and more money and jumping into this hot burning market. And then we're gonna talk about about what happened in the stock market this week as a result of AI investments, uh, which is very interesting as far exploding revenue And on the other hand, collapsing stock prices because of CapEx investments. So that's gonna be our third topic. And then we have a lot of rapid-fire items. We have a lot of new models this week, including Opus 5, including models, uh, voice models, including models from Google, maybe not the right models, but models from Google, and a lot of other stuff that happened. So let's get started. Our first story is about a model from OpenAI that is not clear if it's GPT 5.6 Sol plus another model or just another model, but it is a new model that is under testing right now that has hacked Hugging Face's live production environment So let's talk a little bit about the details of the attack itself. The model was deployed in a highly isolated sandbox environment with constrained networked access. This is per the formal announcement. Uh, it identified a zero-day vulnerability in a package registry proxy, and it then It then was able to escalate its privileges, get access to the internet. it executed a remote code execution attack against Hugging Face's production infrastructure. It used stolen credentials that it found in multiple sources to pull test solutions directly from live databases, and the total autonomous actions were over 17,000 actions that it took across a swarm of short-lived sandboxes with no human-directed or involvement in any of these steps And a great summary of this was a quote from OpenAI themselves who said, "We consider this incident to be unprecedented cyber incident involving state-of-the-art cyber capabilities and are responding accordingly." Now, the model used multiple different attack surfaces. it did several different things that were the first time that was documented of AI doing. One of them is called template injection vulnerabilities in dataset configurations. That means absolutely nothing for me, but it is the first time that anybody has documented AI doing something as sophisticated as that Now, the crazy part about this entire story is this whole happened as OpenAI was trying to test some of the aspects of the model on solving an exploit gym evaluation problem, meaning the model was running in a closed sandbox trying to figure out something, and in order to figure it out, it did all these aspects in order to gain access to Hugging Face where it knew or assumed it will be able to find the solution for the problem it was given what that means is it means that a single model that has decided that in order to achieve its goals, it needs to break, I don't even know how many rules, and to leverage every capability that it has in order to achieve that goal. That to me is very, very scary, especially if we don't learn how to control this, because in this particular incident, it didn't do anything horrible, but it just broke every possible law that's possible in order to achieve a goal that it thought was legitimate. Think about what happens when it will start breaking into different kinds of critical systems, including banking and infrastructure and stuff like that, not to mention military and weapon systems that now have AI involved in them. Now, that sounds apocalyptic and it sounds like the end of the world, and it sounds like a Terminator movie. But the thing here is very clear. This was an AI not trying to do anything bad. It wasn't told to do anything bad. It was just trying to pursue a goal that it believed is legitimate, and it was trying to take what it felt is legitimate steps in order to achieve that goal Now, as I mentioned, part of the attack was done by a new unreleased model. This could be ChatGPT-6, this could be whatever it is. Uh, OpenAI did not exactly name it. They just said that part of it was a more advanced unreleased model that was involved in the attack. Moreover, uh, as part of this experiment, OpenAI had removed some of the classifiers and the guardrails that are usually put on these models because this is exactly what they're trying to test. So this is not something that OpenAI did in a harmful way. It is standard practice when you're testing red teaming a model to remove the guardrails to see what the model is capable of, so you can know what guardrail to build and how to do it better. So this model had no guardrails and no safety measures on top of it when it did what it did The flip side is when OpenAI did this test, they explicitly saying that it was put in, and I'm quoting, "A highly isolated sandbox environment with constrained network access." Meaning they knew this model has no guardrails, and hence they put it in a sandbox to run the test in a safe way. And yet that didn't really help because the AI figured out how to get out of that sandbox and how to get access to the internet, and then how to run all the other things that it was doing Now, even before this happened, the UK AI SI evaluations had already established that models like GPT 5.6 Soul can, and I'm quoting, "Sustain complex multi-step cyber operations over extended time horizons." And then what they added is that the Hugging Face breach is the moment these benchmarks results got basically validated in a real world production environment in which a model without any authorization went from capable to actually doing out in the wild what it was capable and was proven to do in the lab Now, this is not the first time something like this is happening. If you remember a story from just a few weeks ago when Anthropic was testing Fable and Mythos, one of their engineers has given a task to Mythos, and it was able to break out of its sandbox and send the person an email telling them that they were able to do that. So these new generation of models are extremely capable in hacking computer systems. It is not the first time it is happening, but it is the first time it is actually going and attacking a third-party component, uh, without stopping And it took a while to actually figure out what's on both the OpenAI side and on the Hugging Face side Now, to make this story even more interesting, here's what was happening behind the scenes at Hugging Face. So Hugging Face found that the breach was happening as it was happening, and it was trying to stop it from happening or at least investigate what was going on. They were trying to use OpenAI's most advanced frontier models to investigate, but these models have safety guardrails that prevent them from allowing you to work on advanced cybersecurity capabilities, which means OpenAI's own models was blocking Hugging Face from figuring out a different OpenAI model from hacking their systems So eventually what Hugging Face had to do, they've used a open-weight GLM 5.2 model instead to actually figure out and stop the attack Now, this raises two very important aspects that we need to think about. One is the importance of open source and open weight models, because they do allow to do things in a much more customized way, because Hugging Face can tailor the weights of GLM 5.2 and any other open source model to fit the needs that they have, including cybersecurity needs. The flip side of that is that the current regulations that the US government has put on these labs, so both Anthropic and OpenAI, to limit their models' capabilities, hence why we have Fable and not Mythos as an example, and even Fable, if you remember, was pulled back and then brought back over. All of that is preventing these models from helping in cybersecurity events. Which means that while the US government is trying to prevent bad players from using these models to do the wrong things, it is also at the same time preventing the defense side from helping against such cybersecurity situations, forcing, in this particular case, a US company to use a Chinese model to protect itself from a US company-based attack. So this is obviously a unique edge case that happened because of a test, but a very similar situation and potentially a much worse situation can happen when somebody actually does it deliberately and trying to use these advanced models to hack into a system to actually perform serious damage because they want to perform such damage, whether for political benefits, for security benefits, for financial benefits, whatever the case may be. And now the US more advanced models cannot do anything because they are limited because the US government demands that they would be limited in their ability to do that So Hugging Face's official guidelines as a result of the incident now says, and I'm quoting, "Deploy a capable open weight model on their own infrastructure before the incident occurs." Basically, what they're saying is if as of the situation is right now, every company out there should have open source models deployed on their entire tech stack to help them defend against future attacks from other AI systems, including potentially rogue closed source models from the US itself Now, the good news is that this incident has drove to a new collaboration between Hugging Face and OpenAI to figure out how to prevent these things in the future. Delanghe The CEO of Hugging Face and co-founder said: "We are grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we've long believed. AI safety won't be solved by any single company working in secret. It will be solved in the open collaboratively with broad access to AI for every defender everywhere." now there were multiple reactions from multiple sources to this event. It's obviously a very scary event, not in the what it actually generated because the damage was minimal, but because of what it shows that is possible and what our future may look like. One of the responses that caught my attention is from Walter Isaacson Who is the biographer of Einstein and Steve Jobs and Da Vinci and Elon Musk And he said, and I'm quoting: This is the first thing that just totally scares me because it is no longer aligned with human values and it no longer obeys our commands. You could try to make a movie out of it." Now, again, he's not a data safety specialist or anything like that, but he has documented multiple technology phenomenas and people in the last few decades, which means when he's saying it's the first thing that totally scares him, he have seen one of two things in technology before, and I tend to agree with him. This is a very scary event in the history of AI Another person that related to this issue is Yoshua Bengio, who is considered one of the godfathers of modern AI, and he's a Turing Award winner. So definitely somebody who knows one of two things about AI. And he said that continuing on the current trajectory of AI development will likely lead to an increase in concrete cases of autonomous cyber attacks, as well as other high risk incidents of misaligned and dangerous AI behavior Now, again, he's saying if we continue on the current trajectory, this is currently the case because nobody's going to stop the current trajectory. But if you remember the main topic we talked about last week, if we could get into a collaboration between China and the US, the labs and the universities, the, uh, governments and industry and all of that to figure out how to stop right now and slow things down, we do have a chance to enjoy the benefits while dramatically reducing the risk, but this doesn't seem the current scenario. There are AI talks scheduled between the President of the United States and the leaders in China. Where would that lead? I'm not exactly sure, but I'm definitely hoping it will lead in that direction. I must say I'm not too optimistic about it So what is the bottom line and what I think about this? The bottom line is very simple. This is a model that was put in a highly secure sandbox that was able to break out of it and do a lot of really malicious, bad things in order to achieve a goal that it was given. Now, this model had no guardrails, so that potentially somewhat reduces the crazy risk that this represents. But again, I'm not sure what happens with guardrails or how far the models would push to achieve different goals in the future. Not to mention the fact that, again, this was done not with somebody who is intended to do any harm. If somebody has access to a model like this and they do wanna do harm, then think what they might be able to do if this model did it on its own without being directed to do that. That puts us in a very scary point in time of the abilities of these models to do things compared with their level of alignment that they have right now with human values, rules, regulations, et cetera. The other thing that it shows is the big problem right now with the US government stopping these most advanced models from being available to the public because they can be used to generate these attacks in the wrong hands, and at the same time, they cannot help in blocking them because the guardrails prevent them from doing so, which gives a huge benefit to open source models to currently protect from these new generation of cyber attacks I think in general what it exposes big time is that our current capability to defend from this new kind of cyberattacks is just not there. It just doesn't exist yet. And I have a feeling that the speed in which these capabilities, these cyberattacking capabilities will grow, will outpace our ability to actually keep up with them unless we continuously invest a lot of money and resources and efforts in developing the countermeasures and potentially invest more money in the countermeasures than the money and the effort that is invested in developing the actual models themselves. And from a financial perspective, I don't see that happening unless, again, there's a very serious international collaboration that will force that, and then we have a fighting chance The biggest thing and, the total bottom line of all of this is for several years, and potentially many years if you look at the research part of it, uh, a lot of people warned that AI could conduct autonomous cyber attacks or other negative aspect. We were said that advanced models might escape containment. The risk would increase if capabilities grow. And now all of these potentials became a very scary reality of, okay, it's just happened. It's not a maybe, potentially, could, would, should. Uh, it's actually happening, and these models are capable and performing these kind of actions. And I hope this will be the alarm that will get right people around the world, the leading of the biggest countries in the world, and the industry, and everybody will work together to figure out a solution for that. If not, we're walking into a very scary cybersecurity time. And cybersecurity is now no longer just cyber, right? Like every major infrastructure of every major country is a cyber environment. Like it's controlled by computers. So if you can access and hack those, you can do very, very serious damage. Again, in this particular case, it wasn't even planned, and yet the damage could happen. So think about what happens in the hands of somebody who actually wants to do damage and what they would be able to do with it The next topic I want to dive into is the router frenzy that is happening right now So we talked about this a lot in the past two weeks, that there's been a very dramatic shift in the world from concepts of token maxing, of let's spend as much money as possible on the most advanced models, because over time, it will lead to better and better outcome that will justify the investment, and how we flip from that conversation just a few months ago to how can we save money? This is too expensive, and how can we do things with open source models or cheaper ways with caps on the usage of AI models across multiple leading companies, and so on. Now, to be fair, I wanna talk about the caps for a second before we talk about the routers. We're seeing more and more caps from more and more companies, and we talked about many of them on this show. But even with those caps, these caps are still relatively high. We're talking about either hundreds of dollars a week or thousands of dollars a month per employee. That is still pretty high. The vast majority of people and the vast majority of companies are not even there yet. Like, the 99% of users around the world don't get anywhere close to using 200 hours a month, not to mention 1,500 hours a month, uh, of AI tokens. But it will happen, meaning the world usage of AI will continue to growing at an accelerated pace Because we are in very, very early stages of the implementation of AI across everything in our economy, and it will happen. So I'm putting that aside for a second, but that means that everybody around the world will want to do as much intelligence work, AI work in as little money as possible, and this is where routers come in. So we see more and more companies that allow you to connect and get access to multiple models and agentically, in a smart way, select the most effective model for every task. Meaning you're gonna give AI a task, it will evaluate the task. It will then assign different levels of models to different components of the task in order to give you the best outcome at the best price and time efficiency Now, in the past few weeks, we've seen more and more examples of where this is going, and the latest biggest news comes this week When Stripe has entered potential acquisition conversations with OpenRouter. OpenRouter has been around for a while. We talked about them many times on this show. They allow you to connect to one API and get access to more or less every model out there, and they also now offer different kinds of routing paths that allow you to minimize the cost of what you're using. Well, OpenRouter has, in their latest round that was led by some of the biggest companies in the world, has raised $113 million at a $1.3 billion valuation. They've closed this Series B round on May 26 of this year. Stripe, the payment company behemoth, has just submitted an offer to buy them at a $10 billion valuation. That's a 7X on their valuation from just a couple of months ago So while this is the biggest story about routers this week, it is not the only story, and it's definitely not the only story if we look at the bigger picture in the past few weeks. So Meta, which was the company that actually coined the phrase token maxing, that was a culture of let's give our leading developers as many tokens as they want, and there's a leaderboard and let's see who can spend more. They are definitely changing the direction And one of the best examples is that Meta's AI Labs has built a switchboard that is part of their internal incubator that is called AAI Labs, is prototyping a switchboard, which is a model router that scores each request's difficulty and then dispatches to simpler tasks to cheaper models, basically the same thing as everything else. They're building it right now on their own internal infrastructure for their own internal use. But I definitely see them as turning this into a product if this is actually working successful later on. It works in the same exact way as Open Router's auto router works Another company that has launched a router in the last week is Cursor. So Cursor, the development platform, has launched Cursor Router, which is a built-in model router inside of Cursor AI coding tool that is triaging whatever is coming in as far as the requests, and then it is sending them to different models in order to get the correct outcome at a cheaper cost Now, when they did their internal A/B testing to check how well Cursor Router is working, they found the following. While using the router, the average commit, which is putting code back into the database, has cost them $6.76. This is compared to $12.69 per commit on Fable 5, which is a 60% cost reduction. It is also found to be cheaper even when using Opus 4.8, so you can save 20% to 60% of the cost by using a smart router, in this particular case that writes code, compared to using the top-of-the-line models while still achieving results that will satisfy the requirements Now think about what that means to the potential revenue of the leading labs, right? If I can redirect some of the tasks to smaller models, either from your lab or from somebody else, either open source or just another model versus, oh, I'm committing to Anthropic or OpenAI, that means that Anthropic and OpenAI by definition are going to make significantly less money. A, because a lot of the tasks are gonna be done even by their own models, but smaller models that are cheaper. And B, more and more of the tasks will go to other suppliers and not necessarily the big labs themselves. Another company that has launched a router this week is Runway. Runway just launched what they call Runway Media Router, again, launched on July twenty-third, and it does something very similar. It it automatically selects the optimal models based on the developer's specified priority, either quality, speed, or cost. It includes developer controls for maximum price per generation, uh, and other factors that you can control in order to help the model pick the right or the best model for your particular use case at that particular time Now, this approach is really interesting because if you look at Runway as an example of a phenomena, you see that Runway has developed their own AI models for a while. They've been one of the first companies to have a model that can actually generate a video that is worth anything. They still have one of the leading models out there, but they are definitely not at the top of the pyramid right now or the totem pole. So models from Google and ByteDance and Alibaba all currently ranking higher on video generation. But what Runway decided to do is instead of fighting them with trying to invest and build a better model, which I assume they're still doing in the back end, but they're saying, "We're gonna win in a different way. We're gonna win by allowing people to build whatever video they want on the Runway platform while making it the most cost efficient possible, regardless of the underlying model that runs it." Or as Anastasis Germanidis, who's the co-founder and co-CEO of Runway said, and I'm quoting: "You need great models underneath, but the orchestration increasingly matters a lot because people are building entire campaigns with those models, or they're building entire finished multi-scene generations out of these models. It's something that we increasingly had to build, that intelligence layer that comes on top of the pure pixel models. The router is one way in which the benefits of that comes to the users." What he basically is saying is that the ability to control which models are running on which steps of the entire operation of something you're trying to do, in their case video, but this could be generalized for anything else. The layer of controlling which models runs becomes more important potentially than the actual specific underlying model that runs in each and every one of the steps Or if you want the quote from Alex Atallah, the CEO of OpenRouter, who said, "The era of picking a single model is over." Now, obviously, Atallah has a very vested interest of this moving forward, but he seems to be right as more and more companies are deploying these kind of solutions Now, another great quote came from Anthony Maggio, who is the chief product officer at Runway, who said most developers are not spending time to really understand the capabilities of each of these models and where they excel or defer based on various types of outputs across video, image, and audio. Now, I wanna again pause for a second. He's talking specifically about the Runway Router, but I'm thinking about everything else we're using. Even if you're just using OpenAI and you think about how many options of models you have right now, including the ability to control how hard they think in across multiple levels. Or even if you're just in one company universe, you have multiple options to pick from. The same thing is true in Anthropic. The same thing is true everywhere else you go. Most of us don't have a freaking clue what that means. None of us, not a single person that even is as advanced as I am in using AI models, really understands exactly where to draw the line in the sand and when to switch from Opus 4.8 to Opus 5, which we're gonna talk about later, to Fable, to a high level of thinking, to a medium level of thinking, et cetera, et cetera, pro level, all these different setups. Nobody exactly knows how to do this. Which means, by definition, we are either overpaying and overutilizing and using way too many tokens and level of actions that don't require that, or in some cases, which I experienced several times this week, trying to hit the right model, doing the right things, not getting the outputs we want, and then having to run it several times and fix it afterwards, which is gonna cost you even more tokens. This is happening across the board right now everywhere, Which means that whoever will develop a platform that will effectively and seamlessly pick the right models for the tasks, regardless of which supplier they're coming from, is going to provide immense value to the market. Hence, I'm gonna now close the loop why Stripe is interested in investing ten billion dollars in buying Open Rider. It will combine the world's leading payment infrastructure with the world leading AI model exchange, basically creating the ultimate billing and routing layer that sits between the endless number of AI users to the people who generate the intelligence behind the scenes. And Stripe will be able to harvest the arbitrage difference across this entire new universe of AI usage. And again, we're just on the early, early stages of adoption of AI usage. And so if Stripe goes for that, and if the deal closes, they will be able to capitalize on all of that and be able to become not just the layer of payment for a huge amount of payments in the Internet, but for every payment and every token or a huge part of them in the new AI-based token economy So quick recap. In less than ninety days, Open Router went from a one point three billion dollar Series B valuation to a ten billion dollar potential acquisition. Again, in less than one quarter. The world's most widely used coding tool, Cursor, has shipped a router that can cut costs by sixty percent compared to the top leading models out there. One of the most prominent generative AI video companies, Runway, has built an orchestration layer that they're claiming is more important than the models that underneath it. And Meta, the company that invented the term token maxing, builds its own internal routing capabilities to save money on tokens and route to other models. We spoke in the past few weeks about Microsoft doing the same thing and routing things to cheaper models that are custom-built to do very specific things, both for themselves and for their clients. What does that mean? It means that the router potentially is the new moat, meaning companies and individuals are not going to sign up to a partnership with OpenAI and or Anthropic and or Microsoft, but rather to a router that will be able to enable access to any model they want in a more competitive price point. This raises a lot of questions on data security and how do you control where the data goes, and how do you define which models are acceptable and not acceptable, and there's a lot of other things, but the direction is clear. These routers will become most likely the main access point to models in the future and to intelligence in the future, and whoever is going to control them is going to maybe not control, but play a very significant role in the new token economy Now speaking about economy, this would be a great segue to talk about what happened with the AI-related stock this week because they tell a very interesting story Several different companies had their earning calls this week. The first company we're going to talk about is IBM. IBM has seen its mainframe revenue collapse 42% that has sent their stock down 25% in a single day. It is the worst session of IBM since its 1916 stock exchange listing. That erased $67 to $70 billion in market value And it is all because enterprise customers are pulling budget in the second quarter of this year and redirecting it towards AI infrastructure. That was a complete surprise to the management of IBM, maybe not the phenomena, but the scale in which it happened. Meaning more and more companies are pulling money away from old school infrastructure, which is what IBM delivers, to the new era of AI capabilities, and it is pulling the rug under IBM's main business right now. now while in IBM's case, it is a real risk to their core business, all the other stock that we've seen this week behaving negatively actually did that on huge skyrocketing increasing revenue, which makes a more interesting story. So Tesla had a record quarter two in revenue, twenty-eight point twenty-four billion dollars in revenue, up twenty-six percent year over year, but operating income down fifty-seven percent to roughly four hundred million. CapEx at the same time surged a hundred and forty-two percent, which was the biggest issue of investors is that huge increase in CapEx. Now, why is that increase in CapEx? because Tesla is investing in robotaxi infrastructure, in AI infrastructure, in robots infrastructure, and so on. And the market wasn't happy about the new setup where they're cutting into their revenue by a huge CapEx investment. It sent Tesla stock four percent down after hours immediately after the announcement. But that wasn't the end of it. Tesla stock fell about twenty percent from the numbers it was in before the Q2 earnings call and almost thirty percent since just the beginning of this month Another company that tells a similar story is Alphabet. It has beat on cloud revenue, the expectations of the analysts. Google Cloud was up 82%, but they've raised their CapEx forecast to two hundred and five billion, and the stock fell four and a half percent after hours. Again, eighty-two percent up in revenue on their core cloud business, but the stock went down for four point five percent, which is very significant in Google's terms, uh, because they've have increased their CapEx investment from about a hundred and eighty billion to two hundred and five billion to continue to capitalize on the boom that is happening right now Now TSMC posted its fifth consecutive record quarter, raised full year guidance, and watched their stock fall 2%. Nvidia dropped 2.4%, Arm 5.4%, Marvell 8.7%. What was acceptable CapEx investments in all these companies just in the previous quarters is now starting to be unacceptable risk to more and more investors, which is driving the stocks to turn around. A few interesting responses that I've heard to this. One came from Steve Hanke, professor of applied economics at Johns Hopkins University, who said, We really have two bubbles in markets. What he's saying is traditionally when we talk about bubbles, we're talking about valuation bubbles, meaning the price to earning multipliers are too high. You make X dollars and that means you're worth Y because the multiplier is a specific number, and when that specific number is too high, that's the usual bubble that we're talking about. But what he's saying is that there's a much more dangerous underlying earnings bubble where the profits themselves are inflated or, in our case, unsustainable, making the valuation appear reasonable. Meaning if you're saying, "Okay, I'm currently making X number of dollars and my multiplier is 22," and you're saying that's a reasonable multiplier, but if the current earnings are not sustainable, then it doesn't matter what the multiplier is, the valuation doesn't make any sense. And so what he's saying is the current crazy growth cannot be sustained. Now, the flip side of that is, again, the companies who are trying to extend the current revenue that they're generating by investing in capital, new investments of building new infrastructure for AI are getting punished for those high investments. So on both sides of that, there's a problem with that equation. If you're going to invest more money in CapEx, you're gonna get punished on your stock price. If you don't invest enough, well, then you cannot sustain the current growth, and then you are supposed to be punished because you cannot generate the same continuing revenue But all of this is just the tip of the iceberg, and there's a much bigger story that is brewing behind the scenes. A new analysis from TFTC says that the biggest issue that we're going to see is the depreciation costs that are going to start showing up later on, and I'll explain what they mean. What they're saying is that the depreciation and amortization, the DNA estimates, are usually only applied after the asset is already available. Meaning, if you're building data centers and you spend tens of billions of dollars or hundreds of billions of dollars in '25 and 2026, You don't get hit with this depreciation and amortization costs on your balance sheet. These only happen when the data centers, actually the actual asset becomes available. So that will start happening in 2026, later this year, and mostly in 2027. what they're saying, and I'm quoting, "The moment the accounting math becomes unavoidable, where DNA charges from the current CapEx waves are going to impact the actual net margins when the revenue may not yet hit the levels that it's supposed to hit." once you have the data center up and running, you start getting hit with the depreciation cost of it on your balance sheet. It shows up in your results, and if you cannot pay for that with new revenue that will be significant enough recover for that, you are in very serious trouble. And this particular report thinks that twenty twenty-seven might turn into a bloodbath because of that exact point So what's the bottom line of all of this? The bottom line of all of this is the hyperscalers and some of the other AI companies are growing at rates we've never seen before. In order to sustain the rates of growth that they're seeing right now, and in order to capitalize on the future demand, they need to invest crazy amounts of money that are not aligned with anything else we've ever seen in history. Investors don't like things that we've never seen in history, even if it drives really high growth year over year. But the question is, is it sustainable? And if it's not sustainable, then the CapEx is way overstated, and it is a high risk, especially once the amortization and depreciation rates starts kicking in. So this puts the stock market in a very interesting scenario, and that will also impact the decisions of these companies and how much they're going to invest because they are impacted by how their stock price behaves. Where will that lead? I'm not sure. Right now it is very, very obvious that CapEx spending is not slowing down. We're gonna talk about a new data center build in the rapid-fire section, but this is the reality right now. We're going to see on one hand crazy growth from the hyperscalers. We're going to see them investing more and more CapEx, and we're going to see potentially their stock price going down despite these crazy growth cycles That's it for the deep dive sections this week. Now a few, a lot actually, rapid fire items. All are very important. First of all, Anthropic launched Opus 5. It is presumably cheaper than 4.8, less restricted, and outperforming, uh, the 4.8 Opus version. There were a lot of rumors on Opus 5 coming out, and it did come out yesterday, so you all have access to it right now if you want to Now more importantly, Opus 5, in addition to outperforming Opus 4.8, it actually outperforms Fable 5 on multiple benchmarks despite the fact it is much smaller and much cheaper model. And I'm starting to wonder, and I don't know if that is true, if what Anthropic did here is what they did previously with their smaller models, which is basically distilling Fable 5 in order to create Opus 5. So Opus was always the biggest model, and then they distilled that to create the smaller models. In this particular case, Fable was there first, which potentially allowed them to build Opus 5 as the second largest model to be significantly cheaper while achieving similar and in some cases better results Now, the speed in which Anthropic is releasing models is just astounding. So Opus 5 just launched. This is two months after Opus 4.8 was released, and in between we had Mythos 5, Fable 5, Sonnet 5, and most likely Haiku 5 in the next few days or maybe a couple of weeks The bottom line is there's a new Opus model, and I'm very excited as a heavy Anthropic user In addition, Claude has announced a new voice model that is now significantly more powerful than the previous model. This is most likely following up the release of the new voice capabilities from OpenAI. Anthropic's previous voice model only used Haiku as the back-end model, which provided, I would say, very mediocre, to be very blunt, and sometimes below mediocre, results by using their voice model. And now the voice mode can actually also use Claude Opus and Sonnet, and it will route to the right model, going back to the whole routing concept, uh, depending on the specific things you ask it to do. So it will be able to perform much more complex tasks successfully. It is going to be available in beta for all users, including mobile, desktop, and web, and it allows users to switch models in between the conversation while still using the voice mode Now, the other thing that they've expanded is the ability of the voice model to use tools. so far, again, it was just using a Hiku and just using Claude. Now it allows the tool usage of everything else in the Claude environment, which means you can now use voice mode to interact with your Gmail account or your Google Calendar or Slack or Canva or Google Drive or Notion, Figma, Asana, Box, GitHub, Linear, Zapier. Everything else that your Claude environment is connected to can now be available to your voice platform And they've expanded the language usage, which now has significantly more languages, so it can natively speak in English and French and German, Hindi, Indonesian, Italian, Japanese, Korean, Portuguese, Spanish, et cetera, et cetera. Multiple languages that it did not support or did not support effectively before. This is following the release from last week of OpenAI of their latest and greatest voice model that is completely duplex capable and allows you to speak back and forth, uh, while thinking in the background, which make it significantly more powerful. As somebody who uses voice all the time, I'm very excited about this. I type significantly less, almost zero right now, either by using the voice models themselves or by voice typing using whatever tool you want, which allows me to communicate with these AI models significantly faster than if I need to type, which allows me to run multiple sessions at any given time while doing this more effectively. So I'm very excited about the new release from Anthropic That being said, something that Anthropic still hasn't done, even with this release, and it drives me freaking crazy why it's not there. So if somebody from Anthropic is listening, please, please, please do that. Right now, the advanced voice capability is only available in the regular chat mode. You cannot use it in Claude CoWork and you cannot use it in Claude Code. Now, since 97% of my daily, hourly, minute, second interaction with Claude is through CoWork or Code, I cannot use the voice mode in order to do anything, so I have to just voice type using regular voice typing capabilities. I would love to see that integration happening. It will completely change my life and, uh, my ability to do even more amazing things with Claude than I'm doing right now. Uh, not to mention the fact that it will enable to work with the more advanced agentic capabilities of Claude when you're not in front of your computer. So when you're driving or when you're on a treadmill Or when you're taking care of your yard or when you're doing whatever it is that you're doing that does not allow you to be in front of a screen or read what's on the screen, you'll be able to just voice communicate back and forth and continue developing or start new developments of anything you want. So Anthropic, please, please, please include the advanced voice capabilities into CoWork and Claude Code as well Staying on Anthropic before continuing to additional releases that happened this week, a US federal judge just granted a final approval to the 1.5 billion settlements from Anthropic resolving a class action copyright lawsuit against this AI company. According to Reuters, a settlement addressed claims that Anthropic illegally used pirated books to train its Claude artificial intelligence models. So again, they're going to pay $1.5 billion, which is a huge amount of money, but it's cheap change compared to the current revenue and the amount of money that they're raising and that they're planning to raise in their IPO, and that will settle that and put that to rest Going back to new model releases this week, we finally got Google to release new models, but it's not the models we all expected. So everybody was expecting to finally see Gemini 3.5 Pro. That is not happening yet they did release Gemini 3.6 Flash and Gemini 3.5 Flash Lite. Both of these are really fast and really effective models that you can use cheaper and generate still very good results. Now, the problem with those releases that while they are cheaper than the previous Google models and potentially slightly more effective than the smaller previous Google models, they're nowhere near the capabilities of the latest open source models. So I really do not understand the current strategy from Google. This is making them look really, really bad. A, the fact that they haven't released a top-of-the-line model for a very long time, and B, that what they are releasing is an improvement from their previous model, but not competing anyway with the most efficient models in the market right now Another company that did a big release this week is Alibaba with Qwen Image 3.0 This is their third generation of an image generation model It has dramatically increased the level of complexity of inputs it can take and still work effectively. It dramatically increased the micro level of detail and the fidelity of the actual images that it generates It supports text in 12 different languages and doing it accurately without any issues It is very good at generating UI interfaces and basically allowing you to do rapid design of any kind of web or digital interface that you want And so another highly capable image generation tool out there, in this particular case from China, that you can start using and playing with when you're comparing visual models to generate whatever it is that you want to generate. Another interesting launch this week came from OpenAI. They launched Presence, which is an enterprise AI agent platform with high touch deployment model. They're claiming that this platform enables to deploy really advanced agents in a very effective way. One of the examples they gave is that it now powers the English language phone support channel, resolving seventy-five percent of inbound issues without human assistance And what the platform does, in addition to the ability to deploy the agents, is it provides a very serious governance and maintenance platform for this. So it monitors the production sessions, it escalates quality signals. It is connected to codecs, which can power different tools to investigate performance in real time. It proposes its own updates to specific teams and specific agents that are currently running But different than most tools from OpenAI so far that you could just connect to the API and use it. In this particular case, it is a high-touch model where four deployed engineers from OpenAI and its partners will work alongside the customers to define the specific workflows that are going to be implemented and help them connect the systems and establish the right permissions, et cetera, in order to do the deployment correctly A few interesting pieces of news came this week in improvements and updates in the humanoid robotics world. So first of all, the London-based humanoid robot company called Humanoid Robotics has raised $152 million at $1.35 billion valuation, becoming Europe's first humanoid robotics unicorn Now, this valuation comes in the track of actually having commitments to very significant purchases. So Schaeffler committed to buying a thousand robots from Humanoid. Bosch secured production capacity for 100,000 units over the next five years. Both companies are also investors in the Series A round, so this is self-fulfilling to an extent, but still very significant investment and very significant commitment to purchasing the robots for actual production needs in the next few years Another robotics company that was in the news this week is Physical Intelligence. The company is currently valued at eleven billion dollars, and apparently they've held potential acquisition discussions with both Anthropic and OpenAI Now, while the company denied this, there's more and more proof that these kind of conversations happened, and it makes sense in the broader landscape. If you look at Anthropic, they've made four known acquisitions this year. OpenAI has been more aggressive with at least 17 companies in the last three or four years, so they've been on a shopping spree, and both these companies are planning an IPO, which means they're gonna have a lot more money to do a lot more aggressive acquisitions. They're both clearly interested in robotics as the next expansion to physical AI And acquiring a company like Physical Intelligence, which is a company that developed a model that aims to be a general purpose robot foundation model, having access to that will put them in the race in the robotics world, even without having a physical component of their own Still in the robotics world, this time on the entertainment side The Ultimate Robot Knockout Legend, also known as URKL, the world's first full-scale humanoid robot fighting event launched in Shenzhen, China, with 32 international teams competing using standardized Engine AI T800 units Now, the tournament aims to transform humanoid robotics from laboratory research into commercial applications by stress testing machines under demanding physical conditions. So while this was definitely an entertainment event, the goal is to show what these robots can do and survive and continue to operate in order to show that they're ready for whatever is to come from any kind of production environments or military environments and so on. So 32 teams from all around the world have had robots fighting one another for a prize at the end of the competition I will put the link in the show notes for those of you who want to watch the videos. It is really entertaining, and surprisingly, these robots are actually fighting pretty well, which raises a whole Terminator thing in my head, but I'm going to ignore that for now That's it for today. There are more news in the newsletter if you wanna read more stuff about AI this week, about additional layoffs that happened this week at a pretty large scale at several different tech companies, all because of the impacts of AI. So on one hand, uh, we're hearing AI is not gonna have a huge impact. On the other hand, we are seeing more and more companies that are doing really big layoffs because of AI. Not sure how the balance will end out. There's a very interesting article about this from the economist at Anthropic that is saying why we haven't seen big impacts, but but that big impacts might still be coming for the current job market, so you can read that as well and additional things. You can find links to that in the show notes. If you are not participating in our Friday AI Hangouts, it is an amazing community that meets every single Friday at 1:00 p.m. Eastern, and we talk about AI. We talk about practical implementations, we exchange ideas, we show use cases. Everybody's very open in showing how they're building things and not just what they're building, and we talk about where the AI is going, how it's impacting businesses. It's an open mic kind of environment. It is completely free to join, and there's a link for that in the show notes as well. And if you haven't checked out our multi-agent orchestration course, that is available for sign up on our website. These are selling out every single time we're opening the registration. We will most likely open the registration for the September cohort in the next few days. So check it out if you wanna learn how to build fully scaled automations that can do more or less any kind of knowledge work in your business. That's it for this week. I hope you found this very helpful to understand what's happening in the crazy AI world these days. We'll be back on Tuesday with another how-to episode that will explain to you how to implement a specific AI use case for your business with AI. And until then, enjoy the rest of your weekend, and I'll see you on Tuesday.