Hello and welcome to a weekend news episode of the Leveraging AI podcast, a podcast that shares practical, ethical ways to leverage AI to improve efficiency, grow your business, and advance your career. This is Isar Meitis, your host, and what a week we had. I think this is the craziest release week we ever had, and especially if you look at the last two weeks combined, Where we got Fable 5 back, and we got the new GPT models, and we got Grok, and we got lots of new features in Anthropic. In general, like million things were released this week. So the first segment is going to be about everything that was released this week and what it means. Then we're going to talk about potentially the future financial side of this, continuing the topic we had last week about the U-turn in approach from going to, "Oh, let's go token maxing," to, Holy crap, this is really expensive and we gotta cut back and figure out a way to make this financially viable." So there's more information about this, and we are going to talk about the workforce realignment. Basically, the new types of roles that are popping up because of AI in different areas, especially in development worlds, meaning running code or creating code. But that will, I think, impact the rest of us as well. So we're gonna talk about that. And then there's a few interesting things in the rapid fire section as well. So let's get started. so first of all, let's do a quick count of what was actually released. OpenAI released to the public the models they were planning to release last week and the government held them back. So we got GPT 5.6, all three models available to the public, so you can all use it right now. Meta and SpaceX also released models within 72 hours of the OpenAI release So Meta released Muse Spark 1.1 and SpaceX AI, this is now the new name, SpaceX AI released Grok 4.5 In addition, OpenAI released a new voice model that is absolutely mind-blowing. Anthropic released a bunch of new features, and some other companies released additional capabilities as well, all of that within the last 10 days. So let's start diving into a little more details of all these things. Now, the big maybe aspect of this that ties back to continuing the discussion about cost and how that is shifting, maybe the most interesting thing that happened this week is the shift in the entry-level pricing model that became available that was not available before that. So you can now get really advanced AI capabilities at now much more reasonable pricing. So let's start with SpaceX AI with Grok 4.5. It is currently priced at $2 per million input tokens and $6 for input/output tokens. Let's compare that to Anthropic Opus 4.7. So we are two generations back, so we now have Opus 4.8, and we have Fable, which is even more expensive. But Opus 4.7 is currently priced at $5 per input token and 25 output tokens. OpenAI Sol, their smallest model, is now $5 input token and $30 output token. Now, Grok 4.5 matches OpenAI's budget Luna tier on price while claiming Opus class capability. They're claiming the model is significantly faster and twice more token efficient than all the competitors. Or as Elon put it, and I'm quoting, Based on strong positive feedback from customers in our beta test programs at SpaceX AI, we'll make Grok 4.5 available to the public tomorrow. It is an Opus-class model, but faster, more token-efficient, and lower cost." And you're going to see that focus moving forward from everybody because as we mentioned, both the practical aspect of this as well as the trend in the industry is to pull back on the expenses, which mean these companies will have to figure out ways to make the tokens more efficient and will have to deliver higher and higher intelligence at a lower and lower cost in order to stay competitive and potentially in order to stay in business So that was July 8th. On July 9th, Meta launched Muse Spark 1.1, which is priced at $1.25 for a million input tokens and 4.25 output tokens, which is even cheaper than the SpaceX AI Grok 4.5 And it includes $20 free credits when you sign up to use it. Now, it is currently only available in the US. They are claiming competitive performance with GPT 5.5 and Opus 4.8, so not GPT 5.6 and not Fable, but the immediate models right behind that, that as I mentioned multiple times on this podcast, are more than good enough for more or less any task you need to do at work other than really complex advanced capabilities. Zuckerberg describes Muse 1.1 as, and I'm quoting, "Strong agentic encoding model at a very low price," and said that Meta's focus is on, and I'm quoting again, Delivering strong agentic and multimodal models at a very low cost." Again, you can see a very clear focus here from both these companies and the fact that ChatGPT's launch includes three different levels of models with a small one that is significantly cheaper but provides close to the same kind of capabilities shows you exactly where the world is going. The focus is shifting very quickly away from the best performing gigantic behemoth models to smaller, more efficient models. Now, to put things in perspective why I'm saying this, Anthropic Fable 5, which is still available, more on that in a second, is priced at $10 for input tokens and $50 for output tokens. This is five to eight times more expensive than the other models that we just mentioned that are, yes, not as capable, but as I mentioned, capable enough to probably do most tasks that you have Now, between the leading models, between GPT 5.6 Sol and Fable 5, most independent reviewers online currently agree that Fable 5 is more capable. I must admit from testing it that it does some magical, amazing, incredible things, Fable, but it also does some really weird and also annoying stuff that I don't necessarily like. Now, these could be things that will be tweaked in the near future, but as of right now, I think from a capability perspective, it is the most capable model. If you haven't tried it yet and you do have a Claude subscription, you can still use Fable until July 12th. So it was supposed to expire on July 7th, and they extended it to July 12th. I am currently running it across multiple different use cases. I mean Fable 5, just to be safe you understand what I'm saying. Uh, and I'm finding it to be very, very helpful. I seriously doubt it that I will continue using it once it goes to the new model, meaning it's not included in your subscription and it is going to be paid by the token, which is again going to be $50 for a million output tokens. This might be A very expensive fee to pay for something that is better, but I'm not sure it is worth the difference. That being said, what I've done, and I'm recommending to everybody who listens to this on Saturday, you still have a day and a half to use it in your subscription. I upgraded back from the $100 a month subscription I had in Claude back to the $200 a month plan that I had before. So I have more tokens because I reached the limit previously, and I had Fable itself identify out of the 37 things that I'm working on right now, and that's the actual number, I'm not making it up. There are 37 different open projects that I'm working on with AI right now. There are different automations, different agents, different things that I'm doing across multiple aspects of my business and for my clients. I asked it to map the ones that are the most complex and have the highest ROI. So think about a quadrant like the top right quadrant are the things that are the most complex and provide the highest ROI, and we created a plan on how to tackle all of them before it expires. And I'm literally running six or seven sessions in parallel at any given moment, trying to maximize my use of Fable before it goes away. It's gonna get me a lot forward in those short amount of hours that I have before it expires, but this is what I'm recommending to you. Another thing that I've learned that I highly recommend to you to do as well, I asked Fable if it can assign different tasks to different levels of models inside of Claude. And the answer is, it cannot change the main role of it as the leading model inside a conversation. But when it assigns subtasks, it can assign it to any model that it wants. And when it does it in Fable, it's actually pretty cool. It opens these small boxes on the screen, one next to the other, showing you that it is running things in parallel, and then it can assign it to other models. So what I did, I had it go through all the tasks that are open and define the right model for the different tasks. And now when it's running different things, it's not all running Fable. So I'm optimizing my tokens in real time while Fable is actually assigning the subtasks to Sonnet and Haiku and Opus. The interesting thing is, none of the subtasks was actually assigned to Fable itself. So it is running the show, but the actual tasks are done by the other models, and I think this is what these companies should provide in the future. One thin layer of really, really advanced AI model that can figure out the process, the plan, and so on, but then assign on its own to smaller tasks. But like I said, you can rig this on your own just by asking Fable to do this, and I'm sure you can do the same thing with Opus if you don't make it in time for the window of Opus, and it expires and you don't wanna pay the hefty fee. But summarizing the whole pricing aspect of this, in one week, the Frontier class AI tokens dropped from $25 to $50 for a million tokens, depending input or output, to a range of $4.50 to $8 for a million tokens. That is a 75% decrease in cost. And yes, you may not be getting the latest and greatest, but you're getting something that is very, very close to that, that can do, like I said, more or less any job. So am I going to test out the new models from Groq and the cheapest model from ChatGPT and the Muse model? 100%, because I think we can count on US models now without going to Chinese models while getting a lot of intelligence for a lot less money Now, the second important aspect of this that has been very clear in the past few months and is getting even clearer with these latest releases is that the raw intelligence and the raw benchmarks are basically becoming irrelevant. The new competition is fought on something that is on a different axis, and this is how much the model can actually run in the agentic world, meaning how much it can run tools, how long it can run autonomously, how much it can execute multi-step complex work while keeping everything coherent and providing high quality results at a best price possible. All the major launches in the past couple of weeks has been focused mostly around that So as an example, Meta Muse Spark, which is to me the most impressive one, and the reason it's the most impressive is if you think about the new department in Meta, it was set up just over a year ago. There's been a huge mess in there that we reported about many times during the first quarter of its setup and then even longer, and we haven't really seen anything impressive from them. And now they're coming up with the first variation of their initial model. So at release, Meta Muse Spark 1, but now 1.1 is a close to the top of the line at a very competitive pricing. It has a one million tokens context window. It is really good at managing memory across extended sessions. It compacts really well intelligently to allow you to run even longer sessions. It preserves critical steps downstream tasks between different sessions, and it functions both as an orchestrator for planning and delegating parallel, as well as running the sub-agents, all in a highly efficient capability Amjad Masad, who is the CEO of Replit, called it, "What's most impressive about the Muse Spark is how much it packs into one model. Massive million token context, full multimodal support, images, videos, PDFs, built-in search with citations, strong reasoning, top-tier coding abilities, particularly front-end and design, structured output and parallel tool calling, all in a clean, OpenAI-compatible package, a complete agentic foundation. So again, this is coming from the CEO of Replit. He knows one or two things about how AI should run and how it should be specifically in the coding space, and Meta has delivered something that is very competitive in this space. So kudos to Meta, and it will be very interesting to see what the big labs do about it SpaceX AI explicitly positions Grok 4.5 for routine knowledge work, coding, app building, research, writing, and office automation, Excel, PowerPoint, Word, et cetera It is also described as good in handling legal work Or as Elon posted, Our internal assessment is that Grok 4.5 is roughly comparable to Opus 4.7, but much faster. the combination of capability, faster speed, and lower cost is what makes it competitive, and I strongly agree. I do not think we need the most advanced models. You can probably run two, three generations back and still do most knowledge work very effectively. And if you can do that at a much cheaper price, then it is very competitive But these weren't even the most exciting releases of the week. To me, the most exciting release of the week is OpenAI releasing ChatGPT Work. We talked about their plan to release a super app. Well, this is the first big shot in that direction. ChatGPT Work is basically coming to compete with Claude Cowork. It is now built into the ChatGPT app. It is now available in the ChatGPT web version as well. There are more things you can do if you install the app. So the app now has Codex and the regular chat and Work all built into one, very similar to the Claude Cowork environment. It is also including scheduled tasks, just like Claude Cowork, that can run multiple times based on whatever schedule or trigger on specific conditions. It is unified the project feature that is now can work across the board on all these different modes, and it is already available on Pro Enterprise Education, and starting to roll out to Plus members. I have a Plus account, and I have access to it. The one thing that the Plus account doesn't show yet, and I was really, really surprised when I saw that, is that it still doesn't have skills. So if you think about Claude Cowork, skills is the heart of Claude Cowork, and you can build and run and orchestrate skills across multiple aspects of what you do. And to me, this is the biggest magical thing. all the different business plans of ChatGPT. So in the Claude environment, skills exist on all the different licensing capabilities, and in ChatGPT, at least as of this morning, it only exists in the business plan. So in the private plans, uh, at least on the Plus plan, the $20 a month plan, you do not get to create and use skills, and that is a big gap that I assume, based on everything I see from OpenAI and their push to practically copy everything that Claude is doing, uh, it will probably become available sometime in the near future Now, since OpenAI have released Work, GPT Work as part of their web interface as well, and not just in their all-in-one desktop app, Anthropic did the same thing. So now you can use Anthropic Claude Cowork inside of the web browser version of Claude, which is something that did not exist until this week. So now on both platforms, you can use agentic advanced capabilities without having to install the desktop app. Claude also made Cowork available on mobile. So far, all you had is Dispatch, which allowed you to kinda like control your computer remotely from your phone. Now, Claude Cowork exists on the mobile app as well. So you now have the capability to run Claude Cowork Cowork across all the different platforms, and the same thing from OpenAI. This is obviously very exciting for us, the users, because we have more options on where and how we want to develop our capabilities. I am planning to do a very interesting episode, potentially even this Tuesday if I have enough time, to show you how you can work with GPT Work and Claude Cowork on the same things at the same time, which I find really exciting. Uh, A, you can test both of them, but B, you can give the relevant tasks that you think that are one model is better than the other and still work in one cohesive and coherent environment, not losing track without having to copy and paste everything back and forth. More on that again sometime in the next two episodes, Tuesday episodes, so either this Tuesday or the coming Tuesday So what is the current status before we continue to some additional things that were released this week? First of all, we all have access to the latest frontier models from both big labs. So GPT 5.6 Sol, the latest and greatest from ChatGPT, and Fable 5 that was somewhat crippled by the issue with the government. So if you try to now ask Fable 5 anything about biology, including helping your kids in biology homework for fifth grade, it is gonna push back, and it says that it cannot do anything in biology. So it, I think, put it in a too tight of a box than they probably should have, but I guess the government forced them to do that. Still a very capable model. Uh, I didn't get a chance to do a thorough comparison of GPT 5.6 Sol, but from everything I read online, including some of the leading voices are saying it's an incredibly capable model that is going to ask you a little more for your opinion compared with Fable 5. The thing about Fable 5 that I've noticed very well, that it will do everything it can in order to push the boundaries and try to get as close as it can to the final outcome in a solid way, including finding workarounds to issues that it's tackling, which is really incredibly impressive. So if you think about an employee that you hire, you start with an initial employee, and they're gonna ask you a lot of questions, and then they learn over time, and over time, they become more and more dependent and finding more and more ways to solve problems on their own instead of going to the higher-ups to help them solve it. This is what Fable 5 feels like. That is not always the right thing to do because sometimes it doesn't necessarily go in the right direction, or it works on something really stupid for a very, very long time, consuming really, really expensive tokens. Now, again, right now it's part of my subscription, not a big deal, even though it's wasting my tokens to other projects that I'm working on. But if that's gonna be at a pay-to-play kind of scenario, this is going to be horrendous. I literally had it doing something that it needed N8N for, and it failed 20 times trying to do the same thing. 20 times. I'm not kidding. And then I stopped and said, "This is not working. Maybe we should try something else." Said, "Oh, yeah, that's a good idea." But that's after 20 minutes and 20 times of trying to do something that didn't work. And so there are disadvantages of running free for a very long time. But overall, I'm very impressed with its ability to find tools on its own, understand what needs to be done, and actually do it on its own and give you a solid final outcome that you can actually use without bothering you in the middle. I assume GPT 5.6 Sol is the same thing. I did not get a chance myself to test it yet. Now going back to GPT 5.6, other people apparently were testing it for a very long time. So as an example, Pietro Sciarrano, who is the CEO of Magic Path AI, said, I've been testing it for months, and without exaggeration, it is the best model I ever used. Fast, smart, genuinely creative." Theo Brown, the CEO of T3 Chat, said, GPT 5.6 Sol is the world leading in computer use. It made me use it 100X more. When we lost access to 5.6, I quickly started to go insane without it." So two things about what I just read. One is independent people who are heavy users of AI are praising GPT 5.6 Sol, which tells us that it is a very impressive model. Again, I didn't test it myself. But to me, the more interesting aspect that hides here in plain sight, going back to what Pietro Sciarrano said, and he said he's been testing it for months. If he has been testing it for months as a beta tester, it means that OpenAI already has a better model right now in-house, because if they've delivered that months ago, they had months to develop something new. So I'm wondering when exactly we're going to get the next model a full duplex architecture voice capability. It is really, really incredible to use. I've already seen a lot of examples online. Uh, so it has three separate models. One is speech to text, to LLM, to text to speech So what else did we get this week? We got GPT Live from OpenAI, which is a full duplex architecture. Let me explain what full duplex is. Full duplex means that it works like you work. It can speak and think and listen at the same time, which means it works significantly faster, and it simultaneously can do all the three things, which means the conversations becomes a lot more natural, and there are less weird pauses in between. Now, if you think about it, previously, the way models worked, most voice models, there was speech to text. So you said something, the AI converted it to text, sent it to the LLM behind the scenes. The LLM thought about it, converted it back to text to speech, and then spoke back to you. That happens very, very fast, so the delay wasn't crazy, but that's what happened every single time. Right now, it is not what's happening. Right now, this new model simultaneously does these two things. So it listens, speaks, and thinks at the same time as independent capabilities. It provides real-time cues in the conversation, such as mm-hmm" or "yeah," or different cues to let you know that it is listening to you, and it is providing real-time small feedback to what it is doing. And it is coming in two different variations: GPT Live One, which is for paid users, and GPT Live One Mini, which is available to free users, and it replaces the advanced voice mode that was the default inside of ChatGPT before. What does that mean? It means that we can have even more natural conversations with AI, either as an input to the AI itself, which I do all the time, or through the API to develop whatever customer service or any other kind of application you want to develop using this highly capable voice infrastructure Now, if you haven't used voice mode before, I highly recommend it. It enables you to have actual conversations and brainstorming sessions with ChatGPT or, by the way, with any other AI model, and it is very intuitive, and it allows you to share your thoughts, ideas, feelings, whatever you need in a much more fluent way, at least for me. Uh, OpenAI shared at their release that they have 150 million people that are currently using GPT voice and dictation features, and GPT Live is just going to replace the model that does that Now, another interesting thing that was released this week, again, on the smaller scale as far as the impact, but very interesting as far as the feature, Anthropic just launched what they call Anthropic Reflect. It appears on the left side menu if you're on the latest version of the desktop app. And what it does is, as the name suggests, it reflects on the work that you're doing with Anthropic. It provides visualization and that describe your user habits and tell you what you worked on, how frequently you worked on it, and so on. It is now available on all the different membership levels, and it does two things. As I said, one aspect of it is if you want a visualization of your conversations, the topics, the usage pattern, the task types, but it's also prompting critical reflection through questions like what is the one thing you want to keep doing yourself or even that Claude can do it faster than you. But it also provides suggestions on how to improve things you've already built. So this thing serves two purposes. One purpose is to show you how you're using Claude, and the other is to get you to use Claude even more by providing useful suggestions on what you should do next and helps you guide yourself through what you want to do on your own or what you want to do with Claude. I find this like a really cool and smart feature that I am definitely going to use regularly, even though I have my own built-in mechanisms that help me decide what to work on next inside of the Claude universe So what is the bottom line? The bottom line of this first segment is that the race is definitely still on, but the race is changing dramatically. While OpenAI and Anthropic are pushing the boundaries in the latest and greatest models, more and more of the focus is going to more efficient models that may not be the tip of the spear, but are pretty close to that. And in that new universe, Meta and SpaceX AI can now compete as well, together with additional models. And we're going to talk more about that in the next segment where we're talking about the pricing aspect of this. But the race is definitely still on. We have seen crazy amount of releases. Many of them are around the harness and how we use it. if you think about the new ChatGPT Work and ChatGPT app or the new capability to now run Claude Cowork inside of the web app and so on, this comes more critically on the way we're going to do work than which actual model is going to run underneath. I am constantly switching inside of the Claude environment between Haiku and Sonnet and Opus, and now, while I still have it, Fable as well, which means this has become the focus for me as well, and I'm very, very small. I'm a company of one. If you think about much larger companies like many of my clients, the math adds up very, very quickly, and learning how to effectively switch to the cheaper models and how to use them is becoming a very big deal. And so using other models, and not necessarily from China, and again, more about that in a minute, then you can do that now with the models from Meta and SpaceX AI Now the second topic I want to discuss is the financial aspect of the AI race. We talked about this a lot, uh, in the past two weeks. Just a few quick reminders Lindy, as an example, has moved all their traffic from Anthropic Claude to DeepSeek. And Flo Crivello, their CEO, told CNBC, We did it, and you could see the cost curve go down like crash to the ground." DeepSeek is probably less than 10% of the cost of Anthropic. So if you're doing it at scale in a company, that adds up very quickly. We mentioned multiple times, uh, here in the podcast that Uber has exhausted their entire AI budget in just the first four months and then has put a limit to how much AI each of their engineers and employees can use OpenRouter, which is a company that provides multiple AI capabilities through one AI connection. So you connect one API key to OpenRouter, and then you can choose whichever models you want in the back end. I use them for multiple different things, mostly experimentation. They have learned and shared with the world that Chinese models token usage by US companies was 4.5% in early 2025 And is now above 30% since February 8th of 2026, peaking at 46% cost advantage and 60 to 90% cheaper than the leading Anthropic and OpenAI offering. So the reasoning is very, very clear. If you are running the frontier, you are paying huge amount of money that you may not have to pay because you may have cheaper option. And again, in this particular case we see a very clear move in US. So US-based companies that are moving to use Chinese-based open source models Now, is this for every company? Do I think that regulated or security or stuff like that industries will switch to Chinese models? Probably not in the near future. And again, maybe they won't have an option in the near future. We're gonna talk about that in a second. But the other companies, it is very, very tempting Another great example of the usage of Chinese models, we mentioned here before that GLM 5.2 is an incredible coding model. And just to prove that, since it's released in June of 2026, they saw the daily token volume grow 27X and customer account grow 80X in just the first week of its release. That's the fastest module adoption on Vercel tracked in all of 2026 Harpreet Arora, the head of agentic infrastructure at Vercel, told CNBC, and I'm quoting, Price is doing the work here. When a task doesn't need the best model, teams are beginning to route it into the cheapest one good enough, and the recent wave of models coming out of China is winning that trade." And I agree 100%. So the picture is very clear. The token maxing era, which was defined by companies basically telling their employees to max out their tokens, is practically dead. Now, are there still maybe companies that do it? Maybe. I doubt it, but it is very clear that companies that want to scale AI are going to implement some kind of model routing capability that will route them to the cheapest model possible that is still good enough to do the work consistently and effectively. And that is not easy to do right now, but I have a feeling, and I shared that with you last week, that in the immediate future it will become easy through third-party platforms that are already popping up, as well as through the labs themselves. Otherwise, they will not be able to compete Now, if you don't believe me, you should believe Amazon's CTO, Werner Vogels. Now, again, Amazon, one of the largest hosting platforms on the planet. He knows one or two things about how companies are using AI, and he publicly endorsed enterprise shift towards cheaper open source AI on an interview with Fortune on July 10th And he said about frontier models, and I'm quoting, Come with significantly higher operating costs, particularly when deployed at scale." Now, in addition to the fact that obviously he's exposed to exactly how companies are using it because they're hosting all these platforms on AWS and you can get access to all of them, it is even more critical to hear him say that because Amazon has invested tens of billions of dollars in OpenAI and Anthropic so when somebody like that, where your company have over $100 billion in investment in closed source companies is saying that the shift is very clear to cheaper Chinese open source models, it tells you that the shift is very real. So the trend is clear. You have more and more companies are shifting to open source models that are significantly cheaper and are close in performance. In some cases, again, GLM 5.2 is on many of the benchmarks 1% off from GPT 5.5 and Opus 4.8. So it is practically the same at a fraction of the cost. However, there is an interesting development that was reported this week. So on July 7th, Reuters reported that potentially the Beijing Ministry of Commerce is in discussions with Alibaba, ByteDance, and Z.AI over potentially blocking overseas access or restricting in different levels overseas access on Chinese most advanced AI models And they're talking about both closed source and open source models. Now, the exact times and the exact implementation is still under discussion, but the potential future of AI models availability around the world, it is very unclear. So we saw the US government blocking international usage of different advanced models, which again, since then was released in one way or another with different limitations and different guardrails. But the fact that the US government is involved and can impose whatever it decides on future models, and they specifically with Fable said that it should be blocked from international users. They allowed to release it to the US public, just, uh, did not have a way to do that, and so they blocked it from everybody. Well, it seems that it is going to work the same way the other way around So which models are currently in discussion with the Chinese government? DeepSeek-R1, Alibaba Qwen, ByteDance Dubao, Z-AI GLM 5.2. And to put things in perspective, these models are currently driving 30 to 46% of Open Router share of models that we just discussed a minute ago Now, it is obviously not the first time the Chinese government intervenes with AI usage or access to Chinese AI technology by the rest of the world. Uh, in April, Beijing ordered Meta to unwind its two billion dollar acquisition of the Chinese startup, Manus. In early June, Beijing tightened the regulation on overseas deals involving Chinese investors, technology, and data that is related, per them, to national security So this next move perfectly aligns with that. And again, the fact that the US is potentially doing the same thing to them will just increase the likelihood of this is happening. Now, what does that mean and what does that leave us? Well, it means that the US companies cannot rely on Chinese open source models because they may get blocked at a simple decision of the Chinese government, and then you may be left with something with a lesser model that may or may not perform the tasks of your company in the same level of efficiency as the model that you have right now. Now, that puts the other open source providers, the Western world open source providers, and specifically the US open source providers, in a very interesting spot. So the one name that I hope was in the back of your head and we're like, How aren't we talking about them in this whole thing?" is Google, right? We did not mention Google at all with any new releases for a very long time, definitely not in a competitive way. So if we think about the cycle with Google, Google had failed time and time again in the beginning of the new AI era after the launch of ChatGPT, even though they are the ones who more or less invented the technology and had the most advanced lab. Took them a very, very long time to come out of this really bad cycle that they were in, come up with the Gemini models that actually did very well in the beginning. They actually took the lead for a very short time, and now it's been quiet for a very, very long time. I am not sure what is happening in Google. I can give you my two opinions of what's happening. Well, three. One is they just can't get their act together. That would really surprise me because I do think Google has all the different ingredients to lead this race from a compute perspective, from a data perspective, from a advanced lab perspective, talent perspective, even though we told you last week they've been losing a little bit of talent. But they do still have a lot of everything, more than probably everybody else. So I don't think that not being able to do this is the problem. I see that there are two potential things could be happening behind the scenes. Those of you who have been following the AI world know that Demis Hassabis cares about one thing and one thing only, helping humanity come out on the better side of it using his brainpower to do this, and everything else is a noise or a disturbance to that goal. If you heard any interview with him, that's his main focus, is how to solve global hunger, how to solve global warming, how to solve any disease out there, how to etc., etc., etc., improve human lives on this planet, and he will do anything he can in order to do this. There might be two very conflicting things inside of Google right now. One is that they are allowing Demis to do what he wants, and this is the reason we are not getting newer, better models from Google because that's not what they're focusing on as far as the research and the investment. Option number two is that they are trying to force Demis to do what they want him to do, which is develop new products that they can sell, but Demis is pushing back, and they have to play this very carefully because if I have a feeling that if Demis feels that Google is a stopper to what he's trying to do versus an enabler to what he's trying to do, he will leave. And if you have Demis as your leading researcher, you don't want to make him leave. And so it might be a compromise between what Google wants to go as far as pushing a new product versus what Demis wants to do, which is more pure research for better of humanity. And so they are just moving slower than everybody else because of that compromise. I don't know any of that. This is just my personal assumptions, uh, but it will be very interesting to see what's happening. But the reason I brought this up is because Google have very capable Gemma open source models. They are probably the most capable open source models in the Western Hemisphere right now, which means they may get a huge value and benefit from continuing to improve their open source models and releasing those and providing those as a cheaper alternative to the leading models from the US if or when the Chinese models gets blocked. Another interesting approach to this, which is not exactly open source but delivers a similar capability, is Microsoft. Microsoft, in their latest release, in their MAI models, their Microsoft AI models, which are their homegrown self-built models that they're building in order to reduce their dependency on OpenAI, now Anthropic as well, they're not really open source, but they do provide a solution that allows any company to train these models on company data, meaning take the model and make it your own without having access to the open weights and so on, which gets you very close to the same outcome, again, at significantly cheaper prices. So very interesting situation right now going on between the push to get cheaper models and potentially having these models that are currently coming from China being blocked, and how will that impact the open source models from the US and other places around the world, I obviously don't know, but I will definitely keep you posted The last component in the deep dive, and then we're gonna go to a few very quick rapid fires, is really interesting, and it comes from Anthropic's Boris Cherny. Those of you who don't know who Boris is, he's the guy that started and still running Claude Code inside of Anthropic, so he definitely knows one or two things about how to build and use AI effectively So he shared this week what he calls the five archetypes of new roles inside of at least the Claude Code team in Anthropic and how he believes the future usage of AI will be done. He's saying that they're going to basically eliminate the older type of job functions in the code generation world. And instead of the current roles that we know, he sees what he calls five archetypes of positions, or not even positions, but capabilities that people need to have when working with code. The first one is the prototyper. This is the person that generates new ideas rapidly. One of the biggest benefits of AI that it delivers right now is the ability, instead of to think what might be the best solution and then go and build a prototype for a while and then test it, you can prototype almost instantly and then test multiple prototypes at the same time, and then based on actual data, decide how to move forward, because it's actually cheaper to do that than the old way of trying to guess and then build one prototype and then see that it doesn't work. So the first archetype is a prototyper. That's the person that has great ideas and can show them and present them as a prototype. The second one is the builder. He is the person that creates the production-grade product. So he takes the prototype that was actually selected to move forward into production and builds the production-grade product. That comes with a lot of under levels of this, of different aspects that needs to happen for something to go from a prototype to production. Those of you who have done this know that there's a gazillion things to take into consideration, but that's that kind of archetype that he mentions. The next one is a sweeper. The sweeper optimizes and cleans the code and the UI. So if you think about now you have a working application that is in production, it could still run significantly better and be optimized by reviewing the code and making small changes and updating it over time, and this is what the sweeper does. The next archetype is the grower. The grower is the one that iterates towards product market fit. So you have an initial product. You are now seeing how people are actually using it, and you keep on tweaking it in order to get a better fit to the market, in order to grow the market share of the platform through its user base, or maybe find new user bases. So the goal of the grower, as the name suggests, is to grow the usage of the solution based on the actual usage in the real world. And then the last one is the maintainer. It is the person that ensures the system reliability, that it is-- uptime is correct, that it doesn't have any bugs, that it works well with new platforms that come out, that changes in operating system don't break it, API standards, uh, new features, and so on, all these kind of things, just making sure the system runs effectively. Now, these are not job descriptions, and they are not mutually exclusive, meaning you can have one person doing two, three, or all of the above, depending on the company, depending on the team, depending on the needs. Again, in my universe, I actually do all of those because I prototype and I build and I sweep and I do all these things even without thinking about it. So I'm glad Boris gave it specific names Now, under Cherney's models, your actual role is if you want the allocation of time to the different things that you're doing, it's not your job title. So an individual could be spending 60% of their week iterating on features that drive engagement is basically a grower for the 60% of the week. But in the other part of the week, they could be doing other things. So it doesn't matter if they're a backend engineer, a full stack engineer, a frontend engineer, a system architect, all of that doesn't matter. What matters is how you're using AI in specific segments of time and how that promotes the overall eventual usage and adoption of the new thing that you're developing Now, the way I look at this after doing this for a while on daily basis for myself and for a huge range of companies that I work with, from very small ones to really large enterprises, I can tell you that I think similar thing will go way beyond writing code, which is what Chernie's coming from. I do think that more and more people will shift their focus to solution developers of different kinds, and I think different people will hold different aspects of what that means. With all the companies that I work with closely, there are AI champions who are developing solutions. These AI champions, some of them turn into a full-time role as AI champions. That's all they do, is they help integrate and develop solutions. The one thing that he didn't talk about that I definitely see as a very important role is an AI integrator, and the integrator is the one that has more technical skills than the average person, even the average, uh, coder potentially, and can provide the technical assets, resources, and access that other people, the, again, the other archetypes include. So the integrator is the one that will give you access to a specific database in a secure way. He's the one that will allow the access to be privileged to specific individuals based on their role and the things they should see, despite the fact it's connected to your ERP and theoretically can see the entire data. The integrator is the one that will build for you an API to whatever other platform you wanna connect to and will solve all the technical minutia that the average person, and in many cases the average code developer, does not know how to do. So that's the only thing out of the top of my mind that is missing in here. I do see other roles that has to do with customer service. That is not a part of that as well. So how do you service customers with AI? So that's a whole other thing. How to do the engagement inside the team together with AI, so that's a whole other archetype. So there's definitely other things that are beyond what Cherney is sharing that is very focused on delivering code and developing software platforms. But I think the concept is real. I think the concept is old job types is something that may fiddle away and disappear or at least become into a big gray area where it's not exactly clear. And what will be clear is that there are different things that you can do with AI, and you spend different segments of your time focusing on different things. I do that all the time, including shifting from strategic uses of AI or if you want working on the business, where my businesses should go and using AI to that, to the tactical aspects of using the business, which is all that Cherney is talking about. So he's completely not looking at the strategic aspect of this, which you can also do with AI. So again, great concept by Cherney. I think each and every one of us needs to think about our own teams, our own company, our own goals, our own strategy and so on, and see what kind of archetypes we can build for ourselves and for other people in our company, and then think about how we actually manage in that universe is a whole other ballgame Now, we saw aspects of this from other companies as well. So as an example, something of things that I shared with you in the past that connect with this very well. Coinbase announced one-person teams. So inside their company, there are single individuals that are combining the capabilities of engineering, design, product responsibilities all at the same time because they can do all these things with AI, and then it falls back to Cherney's model that they do different archetypes of things in different parts of their time. Another great example that we talked about a few episodes ago is, uh, Cloudflare's CEO talked about measurers, which are the middle management, the people who manage and coordinate costs across different solutions. He talked about these can be removed, not because he wants to target them, but because AI now enables the coordination overhead to go away and can potentially break down some siloed org structures because more people can do more things and get access to more stuff We heard the same thing from Jack Dorsey, CEO of Block, who said, "We're already seeing that the intelligence tools we're building, creating, and using, paired with smaller and flatter teams, are enabling a new way of working which fundamentally changes what it means to build and run a company." So in addition to the huge wave of job cuts, especially in the tech space, what we're seeing is we're finally starting to see where the world is shifting to. It's not necessarily just having less people, it's working in a completely different way. I have a lot of people ask me on how I work and what do I create every single day, and the reality is I create very little. Almost every output that my companies create across the board, I don't create. AI creates it. I ideate, I manage, I help brainstorm, I direct, I make decisions on both strategic and tactical aspect of this, but I don't create anything. AI creates everything for me, from code to applications, to automations, to PowerPoint presentations, to financial reporting, to planning strategically. All of these things I do with AI, and AI generates the output. And this will definitely, once you understand how to do this, is a very freeing kind of feeling that you can do anything. And I think this will eventually, and eventually might take two years, three years, five years, but somewhere in that timeframe, everybody will learn how to do that. And the structure way of different roles doing different things that we are all working through right now will change dramatically, and the companies who will figure out that faster will be able to benefit dramatically Now let's go quickly into a few really interesting and important deep dive segments. The first one is Claude released their co-work data, basically how people are using Claude co-work around the world. They anonymously analyzed 1.2 million co-work sessions from May 2026. And they found the following: about 50% of that is used to administrative and connective tasks rather than the core job functions. So now let's switch to a few quick but very interesting and important rapid-fire items. First of all, Anthropic released their analysis of 1.2 million Claude CoWork sessions from May of 2026, looking into how people are actually using AI or how people are specifically using Claude CoWork. And what they found is this: approximately 50% of Claude CoWork usage falls into connective administrative tasks, status reports, spreadsheets, reconciliation, onboarding checklists, and slide decks that span across roles but definitely not constitute anyone's specific core responsibility. So it's just the day-to-day things that we need to do. That connects very well to what I just said about what I'm doing. I'm not doing the work anymore. Claude CoWork is doing this. I don't know, by the way, how much of those 1.2 million are mine, but I definitely play a role in that statistics because I use multiple sessions of Claude CoWork in any given minute of the day. So business process and operations tasks account for 33.4%. which is more than double the second largest category, which is content creation and copywriting at 16.4%. Now, very different than Claude Code that is used mostly for developers and that are building debugging and shipping code, uh, software development represents only 8.7% of Claude Cowork. And I assume that is because that everybody that understands that AI can write code goes to Claude Code or to any of the other vibe coding platforms in order to ship their code. So when I need to create code, I don't do it in Claude Cowork, I do it in Claude Code and sometimes in Cursor and sometimes in Replit and sometimes in Base44, depending in which of my clients I'm working at that particular moment. But I don't do any of it other than sometimes the planning inside of Claude Cowork Another thing that I wanna touch is there's a very interesting report that actually challenges the AI job loss narrative and the report is claiming that high intensity AI adopters growing headcount by 10%. So a report by Ramp and Revelo Labs analyzing data from over 21,500 US-based companies reveals that significant AI investment is directly correlated with workforce growth and not contraction. So what specifically they found is that companies that are demonstrating high intense AI adoption increase their overall headcount by approximately 10.2% in the first two years of deployment Among the high-intensity adopters, entry-level employment rose by a higher than the average of 12%, leading to a 1.15% point increase in their workforce share of entry-level workers. This is exactly the opposite of everything that we've been hearing that entry-level jobs are going away because of AI As an example, Harvard Business Impact and Education Review indicated substantial employment declines 16% for early career workers in high AI exposed fields Now, what they are stating is that the positive impact on employment doesn't happen immediately, but it gradually happens over time, and it's usually starting to see momentum six to 12 months after the initial AI adoption as they require time for integration and so on. So by the time the AI gets deployed and being used, they understand what kind of new roles they need, and then they start hiring people Now, the other interesting thing is that the employment growth was observed across a wide variety of roles, so not just AI engineering. So that includes sales, up 10.3%, marketing and administrative posi-positions, up 7.8%, finance, customer service, up 6.3%, and even managers and officers, up 6.7%. What do I actually think is happening here is something I speculated about but had no information to support. I do think that AI, in the short term, creates an incredible opportunity, meaning companies who will go all in on AI in the near future can run significantly faster than those who won't. Meaning you have the opportunity to grab market share. Meaning you need more salespeople, more customer service people, more administrative people, and so on, because with the same amount of people you had before, you can do 10x, but those 10x inputs cannot be handled by the amount of work you have right now, so you need to increase your manpower by a little bit to capture that 10x. So with a 10% increase in manpower, you can do x more times the revenue and handle more customers and so on. This requires a highly well-coordinated AI effort that really yields the fruits that you're trying to yield. But I think in the long run, this will not hold. And the reason I think in the long run it will not hold is because I think in the end, everybody will learn how to use AI effectively. Yes, there's gonna be companies who use it more effectively, just like anything else. But the gaps that exist right now between the people who know how to do this and those who don't, and again, I see that every single week when I deliver workshops. I just did a large workshop in Chicago this week, which was really, really amazing, and I appreciate the people who put it together. We had about 70 people, and most of them are in the entry stages of using AI, and this is what I see every time I run these workshops. So the vast majority of people and the vast majority of companies still do not know how to use AI effectively. The gains from just what I showed them in this four-hour workshops will save them dozens of hours every single month. Literally quarter or 30% or 50% of entire positions, but they're not gonna be saved. It will just allow these people to do twice the amount of work they did before for the same company, which means they can serve twice the amount of clients. That may require adding people in other aspects of the company to support that effort, and that is going to be true in the immediate which means you need to figure this out because that is an opportunity that will go away Now, since we mentioned Chinese models, an interesting thing that ByteDance shared this week is that they figured out that Bytedance researchers found that AI agent performance follows a highly predictable mathematical curve when learning from real world environments. What they're suggesting is that if you let AI learn how real life works versus develop them in a lab, they can double their capabilities every three months. So the problem from a learning perspective is that the AI world has more or less consumed the available data, digital data that exist. And so the next thing that needs to be done is learning in real life. And what ByteDance developed is what they called EdgeBench, which is a benchmarking suite of 134 ultra long horizon tasks spanning from software engineering, scientific discovery, formal mathematics, professional knowledge work, each requiring 12 plus hours of continuous agent operation. So not human work, but agent operation. The team logged over 38,000 hours of environment interaction testing for five frontier models, Anthropic Claude Opus 4.8, OpenAI's ChatGPT 5.5, GPT 5.4, plus models from Zhipu AI and DeepSeek in China And what the ByteDance researchers argue is that post-deployment learning from rich environments may deserve the same systematic scaling attention as pre-training has received so far. That will be shifting the focus from static pre-training towards continuous on-the-job adaptation. This is something we've heard before, but now we have proof, like actual research that proves that this may be the next frontier of scaling laws that we haven't tackled yet. The thing here that aligns well with what we discussed before, this will be company specific and industry specific, meaning companies will probably gravitate more towards open source or trainable models so they can do these kind of things in a more effective way The last two topics have to do with some significant changes in the leadership of OpenAI. OpenAI's chief officer and AGI deployment leader, Fiji Simo, transitions to a part-time role after her health crisis. So she took a long personal health-related leave, and apparently she's not coming back, or at least not coming back to a full-time job. So the quote is, "Three months ago, I had to go on medical leave after a severe exacerbation of a chronic illness I've lived with for seven years. During this time, it became clear that the road to recovery would be much longer and more complex than I had anticipated, and that I need to focus on it fully So she's stepping down to a partial role. We'll probably hear in the next few days exactly how that's gonna turn out from a new leadership setup. Uh, we don't have any, or I haven't seen any news on that yet. And then the other interesting one is that OpenAI hires an investment banking expert to focus on disrupting Wall Street OpenAI has opened a new position that is investment banking subject matter expert to enhance ChatGPT's ec- capabilities in high-stakes financial transactions, including merger, acquisitions, and fundraising. So we're gonna see more and more of that, of the leading labs focusing on specific areas of the economy, developing capabilities specifically for that. Because as we mentioned multiple times, and we focused a lot in the last two weeks, it is now less about the frontier capabilities of the models, but actually how well and effective they behave in specific real-life environments, and that is going to be the focus moving forward. That is it for today. I hope you found this episode really educational. I really enjoyed making it. I think a lot of interesting things are happening, uh, right now all at the same time, both on the technological capabilities as well as the financial aspects of this and how that may impact the IPOs that are coming and a lot of other things that are happening in the background, what's happening with the government, and so on. So lots of changes every single day. That being said, on the day-to-day, you still need to learn how to use AI effectively. Come and check out our courses and the multi-agent orchestration course. I think, are sold out of the August cohort, which means we're gonna open September for registration. So don't wait and go and sign up for that, and there's a link in the show notes. We'll be back on Tuesday with another how-to episode, as I mentioned, potentially something interesting showing you how to use ChatGPT and Claude at the same time to work on whatever thing you're trying to improve in your business or your personal life. And until then, have a great rest of your weekend.