The Macro AI Podcast
Welcome to "The Macro AI Podcast" - we are your guides through the transformative world of artificial intelligence.
In each episode - we'll explore how AI is reshaping the business landscape, from startups to Fortune 500 companies. Whether you're a seasoned executive, an entrepreneur, or just curious about how AI can supercharge your business, you'll discover actionable insights, hear from industry pioneers, service providers, and learn practical strategies to stay ahead of the curve.
The Macro AI Podcast
Kimi K3 Explained: Open Weights, Open Source, and U.S. AI Rivals
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Kimi K3 is one of the most ambitious AI model launches of 2026—and it could reshape the global competition between Chinese and American AI companies.
In this episode of the Macro AI Podcast, Gary Sloper and Scott Bryan explain who built Kimi K3, how Moonshot AI created a 2.8-trillion-parameter mixture-of-experts model, and why its architecture is designed for long-running coding and agentic work.
Gary and Scott also clarify the frequently misunderstood difference between open-weight and open-source AI. They examine whether businesses will begin hosting models like Kimi K3 themselves, why most companies will still rely on managed infrastructure, and where smaller private models may deliver greater value.
The discussion also compares Kimi K3 with leading American open models from NVIDIA, Google, OpenAI, Meta and IBM. Finally, Gary and Scott address model distillation, data security, deployment costs, geopolitical risk and the questions executives should ask before adopting a Chinese AI model.
Listen for a practical business explanation of what Kimi K3 means for enterprise AI strategy.
Send a Text to the AI Guides on the show!
About your AI Guides
Gary Sloper
https://www.linkedin.com/in/gsloper/
Scott Bryan
https://www.linkedin.com/in/scottjbryan/
Macro AI Website:
https://www.macroaipodcast.com/
Macro AI LinkedIn Page:
https://www.linkedin.com/company/macro-ai-podcast/
Gary's Free AI Readiness Assessment:
https://macronetservices.com/events/the-comprehensive-guide-to-ai-readiness
Scott's Content & Blog
https://www.macronomics.ai/blog
00:00
Welcome to the Macro AI Podcast, where your expert guides Gary Sloper and Scott Bryan navigate the ever-evolving world of artificial intelligence. Step into the future with us as we uncover how AI is revolutionizing the global business landscape from nimble startups to Fortune 500 giants. Whether you're a seasoned executive, an ambitious entrepreneur,
00:27
or simply eager to harness AI's potential, we've got you covered. Expect actionable insights, conversations with industry trailblazers and service providers, and proven strategies to keep you ahead in a world being shaped rapidly by innovation. Gary and Scott are here to decode the complexities of AI and to bring forward ideas that can transform cutting-edge technology into real-world business success.
00:57
So join us, let's explore, learn and lead together. Welcome back to the Macro AI Podcast. I'm Gary Sloper along with my cohost, Scott Brime. A new Chinese artificial intelligence model called Kimi K3 has generated a great deal of attention. probably have read about it in the news or maybe watch some videos on it. Business leaders are asking who built it, how the company developed something this capable.
01:26
And really whether it represents a serious challenge to the major U S players. We're going to talk a little bit about that today. Yeah, exactly. I think they're also trying to understand, you know, what, what it means when Kimi K3 is described as, as open. So some coverage calls it open source. Other coverage calls it open weight. Uh, those, those terms are used, uh, interchangeably, but they don't mean the same thing. And I think that the distinction affects how businesses can deploy the model.
01:56
uh what exactly they can modify and how they use it and how much they truly know about the way it was created. Yeah, good points. And uh Moonshot AI says K3 contains approximately 2.8 trillion parameters so that it can process extremely long contexts. It can understand both text and visual information, for example. It was designed for coding and extended agent-based work, so pretty impressive. Yeah. uh
02:25
It really has produced some pretty impressive early results. ah But the downloadable weights were not released when K3 launched just recently on July 16th, 2026. And Moonshot says they'll become available by July 27th, right around the corner. So we'll see if that happens. But until it does, ah some of the claims about the model's openness are really ahead of the evidence. And I think the news cycles are getting a little bit ahead of that. Yeah.
02:55
Before discussing the model itself, we should really clarify the terminology. The weights are numerical parameters created during the training phase. that they represent the patterns the model, you know, learn from its data in that environment. Right. Yeah. And when a company releases those weights, like KimiK3 said, well, like Moonshot AI said they were going to do, another organization can generally download the model and operate it on infrastructure that it controls.
03:24
depending on the license, might also be able to fine tune or modify the model, which is what a lot of enterprises are now looking into because they know that there's a lot of cost benefit there. ah And that would be an open weight release. Yeah, good point. I would say also, receiving the weights does not necessarily tell you how the model was built. So just keep that in mind. The developer may not release the full training data set, the complete training code or...
03:52
really the methods used to shape the model after its initial training. Yeah. And open source and an open source AI system goes a lot further than that. It should provide meaningful freedom to use the model, to study it, and then modify and share the system. So open source means that it provides enough information about the code, data and training process for outside users to understand how the model was built. Yeah.
04:22
You know, simple way to think about it, it's kind of like this. An open weight model gives you the finished engine and permission to operate it. A more fully open source model gives you much more design, tooling and information needed to understand how the engine was actually built. Yeah, perfect. Yeah, and so as of today, K3's weights in the final license have not been released.
04:48
And I think the most accurate description today is that Moonshot AI, the Chinese company has announced an open weight model and has promised to release the weights by July 27th. So in reality, calling it fully open source is really premature. Yeah. So you're probably asking, so who built Kimi K3? So it was developed by Moonshot AI as you mentioned, Scott, and it's a Beijing based company founded in 2023. It's best known
05:18
Founder is Yang Zlin, an AI researcher who studied at Tsinghua University and later earned his doctorate from Carnegie Mellon University here in the States. ah Yang actually started his first business in China while studying in the States and went back to China after completing his PhD. uh As far as we've been able to find, uh there was no H1B lottery or visa denial issues. So his background also includes research experience connected to Google and Meta, for example.
05:48
Yeah. And so he went back to China and then Yang founded Moonshot with Zhu Jingyu and Wu Yujin, who also came from Xinhua. So although Moonshot is only a few years old, 2023, it's not a small experimental laboratory. Once it was founded, they pretty much immediately attracted some serious Chinese investment, including backing that's associated with Alibaba and Tencent.
06:18
Hmm. And right out of the gate, Moonshot differentiated itself through long context AI. And we've talked about that in the past, you know, just what context means. uh the company wanted, you know, coming to process longer documents and I think, you know, much larger bodies of information than typical chat bots could handle. So it later expanded into things like coding, visual understanding, and autonomous agents. Yeah. Yeah. So it's now multimodal they say. Right. Um,
06:48
And uh so Kimi K2, K2.5, K2.6 were part of that progression toward K3. And K3 really now represents the company's most ambitious attempt to build a model that can work on, like you said, the long, long projects with less human intervention, which is really where these need to go to be truly enterprise ready. Yeah. So let's get into what makes Kimi K3 different.
07:16
So as we talked about the real headline number is 2.8 trillion parameters with a T very large. so parameters are the internal numerical settings a model learns during training. But K3 does not activate all 2.8 trillion parameters from what I've researched. And that's important because all 2.8 trillion parameters
07:44
you know, are not responding to a prompt in that session. So it's, it's kind of interesting. Yeah. And it uses, K3 uses a mixture of experts architecture or a uh MOE. And we've talked about that a few times in other episodes. uh So K3 has hundreds of specialized neural network components, but only a small number are selected for each token. So Moonshot says K3 has a
08:12
896 routed experts and activates 16 at a time. So making it very efficient using that MOE. Yeah. That's crazy. And think of it as a massive consulting company. I think we've used this example for MOE and other shows. Right. You know, the firm may employ hundreds of specialist teams, but every client problem does not require the entire organization. So the work is routed to the specialist most relevant to the issue at hand. Yeah.
08:40
And that gives K3 access to really enormous total capacity without paying the computational cost of activating the network or the entire model all at once. like we mentioned, K3 is also natively multimodal. So it can process text and visual information within the primary model architecture. Yeah, that's particularly useful in software development. The model can generate code.
09:08
observe a screenshot of the resulting application and identify visual problems real time. And now in that scenario, it can then revise the code and inspect the next result. Yeah. And Moonshot calls this uh putting uh vision in the loop. And I think that's an important step beyond a coding assistant that only determines whether the software runs and works. uh So K3 can evaluate whether the visible output
09:37
matches the intended design. Yeah. Yeah. And so you're probably asking how did moonshot build it? um K3 has attracted additional attention because China continues to face restrictions involving access to advanced American AI chips. I think you've probably seen that in the headlines. Moonshot appears to have responded by focusing heavily on efficiency. kind of what we were just talking about, um, not putting all those trillion parameters in play.
10:04
Yeah, I think one of the uh key components to efficiency, essentially the Chinese firms have been forced to become more efficient without access to the premium chips. And so I think one of the key components is what they call Kimi Delta attention. uh So traditional transformer models become increasingly expensive as the amount of context grows. And that creates a problem when an AI agent is working across
10:32
a large software repository or hundreds of business documents. And Kimi Delta attention gives the model a more compressed work of working memory for, you know, so much more of that processing requirements. uh Instead of repeatedly examining the entire context in the most, you know, expensive possible way, the model maintains a more efficient representation of what has already occurred in those sessions. Yeah. And then, uh
11:01
Moonshot also has improved on how K3 retrieves information from earlier layers, uh distributes the work across its experts and operates at a lower numerical precision. So those changes reduce memory use and improve hardware efficiency, which is really what they were kind of forced to do, like you mentioned. Yeah. I think the bigger lesson is that Moonshot didn't overcome hardware limits through one breakthrough, combined sparse activation.
11:29
efficient memory, lower precision and optimize infrastructure. ah And in K3 is as much of an achievement in systems engineering as it is in the model design itself. Yeah, that's impressive. And there are still some unresolved questions about its complete training process and that you've probably heard about claims of distillation of US models. ah But we're not going to get into that in this episode. But the net is that Moonshot has not
11:58
fully disclosed the training data, the post-training methods, or the role that outputs from other models may have played. ah But the key is that there appears to be a real original engineering innovation in this model from the Moonshot team, which is, like we said, it's impressive. But the full technical report, ah it will probably surface some of that detail, assuming it's all in there. If not, people will uncover it all pretty soon.
12:26
Yeah. And the research uh appears that K3 has performed very well in early testing, particularly in coding, visual software development, and extended agent tasks. So some evaluations place it within the same broad capability range as advanced systems from folks like OpenAI and Thropic, which is obviously big news for these U.S. companies that are spending massive amounts of capital. You're probably seeing how the markets have reacted to that. ah it's definitely an interesting headline.
12:55
Yeah, they definitely did it much more efficiently from a capital perspective. And Moonshot acknowledges that the strongest proprietary models will still outperform K3 overall in some areas, but K3's most distinctive strengths seem to be long context running coding and work that combines software development with visual feedback. So they're right in the top spots for that. Right.
13:23
Benchmark results still need context. So scores can be affected by the tools available to the model, the system prompt, the reasoning budget, and the environment in which it takes place, right? So I think you have to keep in mind that a leaderboard does not automatically identify the best model for a particular business workflow. So just keep that in mind. Yeah, right on. And cost also needs to be measured carefully.
13:48
Cost obviously is uh a hot topic now. We've got AI for FinOps and CFOs are starting to really understand this. uh So K3 might offer an attractive price per token, but early testing is starting to show that it can generate a large amount of output while trying to solve a problem. So businesses really need to measure cost per successful result. And so we've been talking about that in other episodes is outcome-based success and cost, not simply
14:17
you know, cost per million tokens. Yeah. The meaningful question is how much it costs to complete the task, the model that AI Phenops is moving to, as you just alluded to, Scott, because a model with inexpensive tokens can still be expensive if it's, you know, using many more of them in those sessions. um just something else to keep in mind. Yeah, exactly. Yeah. And on that note, Moonshot specifically acknowledged some of the early uh
14:46
uh output that that K3 might become overly proactive when instructions are ambiguous. So if you're not telling it precisely what you're looking to do, it's going to spend a lot of tokens trying to figure out what you what you want for output. ah So that that can be helpful in a sandbox, but it becomes a lot more serious when the model has access to financial customer or operational systems. Right. All right. Well, let's let's shift to what K3 means for for
15:15
business, I'm sure a lot of the business leaders are asking that question while you're listening to the episode. ah The availability of open weights does not mean every company will begin operating in its own K3 infrastructure, even though many will probably look to do this. The model is enormous at four bits per parameter. The raw weights alone would require roughly 1.4 terabytes, for example. Yeah. And that does not include context memory. ah
15:45
runtime overhead or the infrastructure needed to support multiple users. So Moonshot put in their recommendations at least 64 accelerators for a serious deployment. So that's not really something that the average midsize enterprise is going to install on an ordinary server. um Most companies would access K3 through a hosted service or a managed private deployment. The model might operate inside the company's cloud account or private
16:13
environment, co-location environment, while a provider manages the underlying hardware for them. So it's likely that a percentage of spend that would have gone to the likes of OpenAI or Anthropic will shift to other models of deployment. Yeah. Yeah. think those companies, if they were public right now, would probably see a pretty big hit with the realization that this type of model deployment is available. you know, a private deployment
16:40
will provide more control over data and model behavior than using Moonshot's public application. ah But it wouldn't eliminate security licensing or governance responsibilities. The company still is going to have to control which systems the agent can access and what actions it can perform if you host it in a private model. And at the macro level, at least initially, ah
17:07
Smaller open models will probably have a greater impact on everyday enterprise self-hosting. So most companies, they don't need 2.8 trillion parameter models to classify documents, extract information, or I don't know, answer questions from an internal knowledge base, for example. Yeah, I think people are starting to realize that the future is a portfolio of models across the enterprise. uh I think we have a...
17:34
an episode teed up about Microsoft strategy in a business might use a frontier proprietary model for its hardest reasoning tasks and a managed open model for sensitive work and small local model for those repetitive high volume operations. So getting smarter about which models you use, how much they cost for your particular workflows and use cases. the strategic change is not that every company will host K3 with
18:04
more open source models, businesses will have more choices about where models run, who controls them and how easily they can be replaced within their own internal ecosystem.
18:21
ah Yeah, so let's just talk a little bit about which American models are going to compete in this K3 space. ah So in the US here, there's a lot of concern about US falling behind with open models, but there really is a substantial and a growing open model ecosystem. And the releases are just going to keep coming. ah But as of right now, there's really no exact American equivalent to something that's the size and power of Kimi K3.
18:50
in the open model space. think the leading US models are generally emphasize efficiency and practical deployment rather than trying to match K3's enormous total parameter count. ah it's a model that on the Chinese side, it's something that they wanted to launch. Right. NVIDIA's NemoTron family is one of the strongest enterprise challengers. Right. uh NVIDIA combines open source models with its hardware and enterprise software ecosystem, which many of you are probably aware.
19:20
Um, it's advantage may not be model size, but the ability to operate efficiently on infrastructure that businesses use already today. Yeah. Yeah. Huge ecosystem there with Nvidia. Um, some, uh, smaller us models from Google and open AI, uh, compete by being much easier and cheaper to deploy. And, uh, we talked about open AI and Profic quickly building their, their armies of forward deployed engineers.
19:47
And I think one of the reasons that they're really focused on that is to help companies quickly get into the game with hosting these models. And then obviously using the large frontier model for certain specific tasks. And then Google, of course, you've got the Gemma family. Open AIs is called GPT-OSS models and they can run on dramatically less hardware than K3. So for many, for a lot of business processes, that practicality may matter more than
20:16
you know, maximum capability, which is what, you know, K3 represents on the Chinese side. Yeah, exactly. Meta's Lama family remains important, um, because of its broad, you know, developer and cloud ecosystem. And Lama may no longer have the uncontested open model capability lead at once held, but it's tooling, integrations, um, the available expertise in the market. Um, yeah, continue to make it relevant. And there's, there's a lot still on their roadmap.
20:46
Yeah. Yeah. I think American developers are generally competing through efficiency, accessibility, and moving to complete enterprise support. And Kimi K3 is making a different statement. So Moonshot is attempting to prove that an open weight model can operate near the frontier of global AI capability. ah And availability of a model like Kimi K3 has huge security implications, which is all over the news today.
21:16
since there are obviously bad actors all over the world that are going to try to access it, support it, use it. And, we'll, we'll cover that in another episode. That's a couple episodes all by itself. Yeah. It's a big topic and the government put a pause on the release of fable five and shortly thereafter, Kimmy, uh, three is something you can download. ah Um, so we'll work on an episode about what that means for security, uh, in the future. Yeah.
21:43
So, so what do business leaders need to know or what you should do now? I think a company considering K3 should begin with a real business workflow and determine if the, existing ecosystem offers solutions that will work for them. It, it should not make a strategic decision based on public leaderboard or dramatic demonstration, for example. Right. Oh yeah. Yeah. The, evaluation should really measure output quality, completion time.
22:11
uh employee corrections and just total cost. uh And early testing should use public synthetic or sanitized information until the company really completes the appropriate level of privacy and security reviews, their due diligence. Yeah. And businesses should also determine whether a smaller model can perform the work. The largest model is not automatically the best business model either. Right.
22:40
Yeah. And I think, think businesses, companies should also avoid becoming, you know, completely dependent on one provider. uh Their business rules, data retrieval, workflow orchestration, those things should be separated from the underlying model, wherever practical. Obviously things are going to change out in the marketplace. They can change very quickly, even with a simple government order. And that'll make it easier to test and replace models as the market or
23:09
whatever changes. The next major milestone for Kimi K3 is July 27th. I think we talked about that at beginning of the show. And this is when Moonshot says it will release the weights and additional technical information around the platform. Yeah. The final license, uh there'll be more independent testing and then you'll be able to understand realistic hosting costs. And that'll tell us a lot more about K3's actual enterprise potential. uh Security risks aside.
23:39
So the launch is obviously established at K3 and all of their technical breakthroughs are certainly important. um But what happens after the weight release will determine how practical it really is for the marketplace, or if it's just more of a statement. Agreed. So to wrap up this episode, Kimi K3 challenges the idea that most capable artificial intelligence models must remain closed.
24:07
It also challenges the belief that American AI labs have an insurmountable lead over everyone else. uh And from a technical perspective, K3 certainly demonstrates how architectural and infrastructure efficiency can compensate, at least in part, for limitations and access to the newest hardware. So that alone has implications for both the global AI industry and global
24:37
competition. Yeah, but K3 does not mean every company should download a Chinese model or build an AI super node. Open weights do not eliminate infrastructure costs, uh licensing uncertainty or security responsibilities. It really becomes a detailed cost benefit analysis and infrastructure decision and ultimately a security decision. Yeah. Just another thought. think uh this Kimi K3 release is kind of a
25:07
a preview of a world in which frontier intelligence, because K3 is really at that frontier level, is available from more countries, more providers, and available for more deployment models than just using Anthropic Clawed or just using OpenAI. So it definitely creates complexity for business leaders, but it also creates leverage and flexibility. Yeah, this is an important topic and one that
25:35
really has huge implications for the AI race. We'll continue to follow the K3 weight release, independent evaluations in the market and in response from American model developers. Thank you for listening to the Macro AI podcast. Please continue to share our episodes with your colleagues and send in your questions. And until next time, look forward to speaking with you.