Yesterday in AI
A rundown of all of the important stories in AI that happened yesterday in 10 minutes or less.
Yesterday in AI
Capability vs. Size: Alibaba’s Qwen3.8, Google Custom Chips, and China's Romance Ban
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Yesterday in AI | 21 July 2026
Capability vs. Size: Alibaba’s Qwen3.8, Google Custom Chips, and China's Romance Ban
The global AI race enters a hyper-accelerated phase this week with massive shifts across software, custom hardware, and social policy. In this episode, we break down Alibaba’s new 2.4 trillion parameter Qwen3.8-Max-Preview, scrutinizing their claim of being the world's second most capable model against benchmark realities. We unpack Google's secret "Frozen v2" server chip project, aiming to hardcode Gemini directly into silicon to combat an unprecedented cloud capacity shortage.
We also follow the money behind the infrastructure crunch, covering Databricks' stunning $188 billion valuation and CuspAI’s $450 million raise to find rare metal substitutes for semiconductor manufacturing. Finally, we explore China's landmark decision to ban AI companion apps to protect falling birth rates, and review new data from Epoch AI proving that style-matched AI text completely breaks modern detectors.
Feedback? Email mike@yesterdayinai.news or connect on LinkedIn, X, or Bluesky. If you like the show, please take a minute to rate and review it so others can find it!
Hi folks and welcome back to another edition of Yesterday in AI, your daily digest of everything happening in the world of AI in roughly 10 minutes. I'm Mike Robinson. It's Tuesday, July 21st, and it's been a wild few days. China unveiled the 2.4 trillion parameter model at the Shanghai World AI Conference. Google reported Monday it's baking its own AI directly into custom silicon, and the same country that just dropped a massive new model with sweeping capability claims and zero benchmarks also became the first to officially ban AI boyfriends. It's that kind of week already. Let's get into it. Alibaba unveiled Quinn 3.8 Max Preview on Sunday, coming in at 2.4 trillion parameters. For context, Moonshot AI's Kimi K3, which we covered last Thursday, has 2.8 trillion parameters. Alibaba isn't making a size flex here. They're making an efficiency and capability play. Their headline claim? Quen 3.8 is the second most capable model in the world, trailing only Anthropics Claude Fable V. That ranking claim deserves real scrutiny. Independent benchmarks from artificial analysis currently show Kimi K3 sitting third globally with a score of 57.1, behind Claude Fable V at 59.9 and GPT 5.6 Sol at 58.9. For Quinn 3.8 to actually claim second place, it would need to jump over GPT-5.6 Sol while running on a smaller overall parameter footprint than Kimi K3. That's certainly possible. Parameter count isn't everything, but right now, there's nothing independent to verify it. Here's why. Alibaba released no benchmark data, no model card, no technical documentation, and no activated parameter count. That last one matters because in a model this size, total parameter count could be a vanity metric. Big AI models often use a mixture of experts design where only a fraction of the network fires on any given query. The total might be 2.4 trillion, but if only 200 billion are actually working, the effective model is much smaller. Alibaba hasn't said, so right now, second in the world is purely their claim with no receipts attached. What we do know, the model handles images, video, and documents alongside text. A preview version is live now on Alibaba Cloud at 110th the normal price, with full open weights promised soon. The broader context is what makes this wild. This is the third major publicly released model in about a week from Chinese labs. DeepSeek V4 Pro is priced at a rock bottom $4 per million tokens. Kimi K3 dropped last week and now Quinn 3.8. They're compressing their release cadence and undercutting US pricing, all while working around strict hardware export controls. If those performance claims hold up once independent testing lands, the geopolitical model race is about to get a lot louder. While Alibaba is busy trying to maximize capability on software models, Google is taking the exact opposite approach, trying to solve the compute crisis by trapping the model inside custom physical silicon. A report out Monday from the information picked up by Reuters says Google is developing a server chip internally called Frozen V2. The concept is to hardcore specific parts of Gemini's architecture directly into the silicon. It's purpose-built circuitry designed to run Gemini and absolutely nothing else, etched in at the hardware level. To be clear, this chip is built for inference, not training. Training is the expensive month-long process of teaching a model. Inference is what happens every time you send a prompt and the AI responds. That's where the massive operational volume is, and that's where hardware efficiency translates directly into saving millions of dollars. The projected gain is massive, six to ten times more output per watt of electricity compared to Google's current chips, with a target deployment of 2028. But there's a glaring risk embedded in this approach. If Gemini's architecture changes significantly before 2028, these chips become multi-million dollar paperweights. You've essentially built a massive highway for a car that might get discontinued next season. But the reason Google is taking that gamble tells you everything about the current hardware crisis. Reuters also reported that Google Cloud has started declining external customer deals because they literally don't have enough compute capacity to support them. When demand outstrips supply so badly that you're turning away paying enterprise clients, you start taking architectural bets you otherwise wouldn't. Frozen V2 is Google's bet that Gemini's core design is stable enough to bake into hardware permanently. Whether you're hard-coding models into custom chips or just trying to survive the current hardware drought, the real winners of this compute crunch aren't the labs building the front-end models, it's the infrastructure companies selling the enterprise pickaxes. Which brings us to Databricks. The company announced a massive new funding round last week at a mind-boggling $188 billion valuation, up from $134 billion just five months ago in February. Because why let a little thing like gravity slow down evaluation? The round is led by COTU, expected to raise roughly $3 billion and will close later this summer. Databricks builds the data platforms and governance tooling that enterprises need to actually deploy AI without getting sued or going bankrupt. The three products getting the new capital are telling. Unity AI Gateway for multi-model cost control, Genie, their AI coworker for internal business data, and Lakebase, a managed database built specifically for AI agents to read from and write to. That framing is entirely deliberate. Databricks has spent years positioning itself as infrastructure, safely removed from the brutal model arms race. And right now, that's where the real money is. Enterprises are discovering that getting a frontier model API key is the easy part. Knowing which data goes where, controlling spiraling costs, and keeping your autonomous agents from going rogue, those are the hard parts. Databricks is betting companies will pay a premium to solve them, and Wall Street agrees. While Databricks is securing multi-billion dollar valuations to manage enterprise software data, another startup just raised a massive war chest to fix the actual physical constraints, holding back the hardware itself. Cusp AI, a Cambridge-based startup, closed a $450 million Series B round, pushing its valuation to $2.6 billion, up from $520 million just nine months ago. Backers include heavyweights like Kleiner Perkins, NEA, Jeff Bezos' Bezos Expeditions, and AMD Ventures. Along with the Cache, they launched the AI Materials Foundry, a coalition of over 45 giants, including Nvidia, Meta, Samsung, Hyundai, and Lam Research. What does Cusp AI actually do? Their Mira platform uses AI to predict how new materials will perform before anyone ever touches a beaker in a physical lab. Instead of spending months running manual experiments on a candidate material, Mira models how it behaves at an atomic level and tells you if it's worth manufacturing. This year, 80% of Cusp AI's RD focus is on finding substitutes for supply-constrained rare metals used in semiconductor fabrication, specifically ruthenium and iridium. Both are incredibly hard to source, wildly expensive, and heavily concentrated in a small handful of countries. That makes this directly relevant to every conversation about chip supply chains and data center constraints. The chip problem isn't just a logistics problem, it's a materials science problem. Cusp AI is betting AI can compress decades of materials discovery into a few years. Bezos and NVIDIA just bet nearly half a billion dollars that they're right. Now, shifting from the physical composition of microchips to the emotional fabric of society, it turns out regulators are starting to realize that AI isn't just threatening the rare metal supply, it's threatening human romance. In a fascinating policy shift, China became the first major country to specifically ban AI companion apps that simulate romantic or family relationships. The new regulations require platforms to build in instant exit options, mandatory regular reminders that the AI isn't a real person, and hard limits on long-term emotional memory, the exact feature that makes an AI companion feel like it actually knows and loves you over time. For users under 18, virtual partners are banned entirely. ByteDance, Alibaba, and Tencent all looked at the compliance checklist and immediately chose to suspend their companion features entirely rather than try to comply, proving that breaking up with millions of users simultaneously is much easier when they're just lines of code. What stands out here is the framing. The official grounds for the ban are China's falling marriage and birth rates. Regulators genuinely believe AI relationships are substituting for human ones. In a US or EU regulatory context, you'd expect the argument to center on data privacy or psychological manipulation. China's argument is much simpler. These products are working too well. People are choosing algorithms over marriage, so the algorithms have to go. While China is trying to force people to look for real-life human partners instead of algorithmic ones, the academic and legal worlds are still struggling to figure out if the words right in front of them were written by a human or very convincing digital ghostwriter. Epic AI published new research this week testing three widely used AI text detectors, Pangrum, GPT-0, and Originality.ai, against large samples of both human writing and AI-generated text. The results are a bit of a wake-up call. When AI uses a basic vanilla prompt to generate text, detectors catch it almost perfectly, with a false negative rate below 1%. But when AI is given samples of a specific author's writing and explicitly asked to mimic their style, the detection numbers completely collapse. On average, detectors miss about 13% of style-matched AI writing. For scientific writing specifically, that number jumps to 26%. Worse yet, originality.ai incorrectly flag 3.8% of actual human writing as AI generated. That's nearly 1 in 25 real human passages getting wrongly accused, which is fantastic news if you enjoy arguing with an automated academic grading system that thinks your syntax sounds suspiciously robotic. The practical takeaway? If someone is using AI while deliberately matching their own voice, current detectors are essentially guessing. This connects directly to that case we covered last week where a federal magistrate in Connecticut ordered an expert witness to hand over her AI prompts as part of her courtroom methodology. The detection problem and the disclosure problem are officially colliding, and the legal and academic systems are completely unprepared for the fallout. And that's it. If you have any feedback about this show, you can email Mike at yesterdayanaai.news, or you can find me on LinkedIn, X or Blue Sky. And if you like this podcast and want to see it continue, please take a minute to rate and review it so others can find it. Thanks. As always, thank you for listening today. Stay curious, and I'll see you tomorrow.