The AI Daily Brief: Artificial Intelligence News and Analysis

Is Kimi K3 Really Fable Class?

28 min
Jul 17, 20263 days ago
Listen to Episode
Summary

The episode analyzes Kimi K3, a 2.8 trillion parameter open-weight model from Moonshot AI, examining whether it truly reaches Fable 5 / GPT-5.6 SOL frontier class performance. Benchmarks show K3 is the strongest open model ever released, placing third on Artificial Analysis's Intelligence Index, but real-world testing reveals gaps in deep engineering tasks, high token costs, and slow inference. The episode contextualizes K3 within the broader US-China AI race, noting the distillation narrative is fading and the gap between Chinese and Western frontier models has narrowed to potentially under three months.

Insights
  • Kimi K3 is the most capable open-weight model ever released, but benchmark performance overstates practical utility — real-world debugging and long-horizon tasks still favor Fable 5 and GPT-5.6 SOL.
  • The 'Chinese models are just distillation' narrative is increasingly untenable; credible researchers and practitioners now acknowledge genuine innovation from Chinese labs.
  • K3's cost and compute requirements ($5.40/M tokens blended, ~44 Mac Studios to self-host) undermine the assumption that open-weight Chinese models are cheap alternatives to Western frontier models.
  • The lack of safety guardrails on K3 raises serious policy questions about how open-weight frontier-class models should be vetted internationally, especially given the US recently restricted Western equivalents on cybersecurity grounds.
  • The trajectory of open-weight model capability is tracking closed-source frontier development, meaning policy frameworks must account for both curves simultaneously.
Trends
Open-weight models are converging with closed-source frontier models in capability, eroding the performance moat of proprietary labs.Chinese AI labs are closing the frontier gap faster than Western analysts predicted, with estimates now suggesting less than 3 months behind leading US models.Open-weight frontier models are approaching closed-model pricing, challenging the assumption that open-source equals cheap.Safety and guardrail gaps in open-weight frontier models are becoming a critical policy and enterprise risk concern.Mixture-of-experts architecture has become the de facto standard for both open and closed frontier models.Local/self-hosted AI is becoming more viable at the high end, but compute costs remain a significant barrier for most organizations.Model benchmarks optimized for visual UI coding tasks are increasingly seen as poor proxies for real-world engineering capability.The distillation debate is giving way to recognition that Chinese labs possess genuine foundational research and training capability.International AI governance frameworks are lagging behind open-weight frontier model releases, creating regulatory uncertainty.Enterprise AI adoption is shifting focus from model access to sophisticated collaboration behaviors, as evidenced by KPMG research cited in the episode.
Companies
Moonshot AI
Developer of Kimi K3, a 2.8T parameter open-weight model claiming near-frontier performance.
Anthropic
Maker of Claude/Fable 5 and Opus 4.8, used as primary benchmark comparison for K3.
OpenAI
Maker of GPT-5.6 SOL and GPT-5.5, key benchmark comparators and competitive reference points.
DeepSeek
Chinese lab whose R1 model triggered major market panic; V4 Pro cited as cost benchmark contrast to K3.
Z AI
Chinese lab whose GLM 5.2 model previously generated frontier hype; K3 scored 6 points ahead of it.
Artificial Analysis
Independent benchmarking firm that confirmed K3's Intelligence Index score of 57, placing it third overall.
Nvidia
Suffered largest single-day dollar market cap loss in stock history following DeepSeek R1 release.
Cognition
AI company whose ambassador Justin Goria praised K3 as a milestone, especially for 3D and front-end tasks.
Vercel
CEO noted K3 is the best-performing model on NextJS.org evals, ahead of all proprietary models.
Arena AI
Ranked K3 number one in front-end code arena, a 17-place jump, and top in six of seven domains.
Mixpanel
Founder Sue Hale cited to support the view that Chinese lab distillation narrative is overexaggerated.
Carnegie Mellon University
Alma mater of Moonshot engineer Jinyu Yang, cited in context of Chinese AI talent pipeline.
Xiaomi
Maker of Mimo V2.5 Pro, a 1T parameter model cited as context for K3's unprecedented 2.8T scale.
Thinking Machines
Released Inkling model just under 1T parameters this week, cited as context for K3's scale milestone.
OpenCode
Developer Dax shared anecdote where K3 used 3x the spend of SOL and failed a simple hover color fix.
People
Guillermo Rauch
Noted K3 is the first open model ahead of all proprietary ones on NextJS.org web engineering benchmark.
Ethan Mollick
Provided nuanced K3 testing results including shader praise, murder mystery failure, and stats audit issues.
Alex Finn
Declared K3 fundamentally changes the AI race, claiming it beats Fable 5 on some benchmarks.
Jinyu Yang
CMU PhD who joined Moonshot; shared insider account of Moonshot's AGI hunger culture driving K3 release.
Nathan Lambert
Called for distillation arguments to end, affirming China is genuinely good at building models.
Shriram Krishnan
Summarized K3 as a big moment with multiple implications for the entire industry.
Simon Willison
Highlighted K3's extreme token consumption — 13K reasoning tokens for a simple SVG benchmark task.
Dan Shipper
Expressed strong skepticism about claims that K3 matches Fable, calling for a proper vibe check.
Jeffrey Emanuel
Found K3 impressively reviewed a 1.5MB plan already exhausted by Fable and SOL, but noted it's not equal.
Elon Musk
Predicted Chinese labs would reach Fable 5 class by Q1, prompting Z AI founder to say it'd be sooner.
Joe Weisenthal
Retweeted code arena results questioning whether K3's rise was behind the NASDAQ dropping.
Justin Goria
Called K3 a new milestone and demonstrated its 3D and front-end capabilities with a Minecraft clone.
Tyler John
Noted K3's bio safeguards are less comprehensive than Fable's and argued Chinese innovation is real.
Vai McCoy
Warned that fine-tuning K3 into a malicious coding agent will be trivial given open weights access.
Sanyam Satya
Internal front-end eval placed K3 closest to Opus 4.7, about three months behind current frontier.
Quotes
"Do not confuse a gorgeous demo with real engineering ability."
Divya (AI Engineer)
"Chinese open source is no longer six months behind, but it's also no longer 10% of the cost either."
Jeff Wang (Cognition)
"The era of the Chinese labs being far behind is over. Kimi is at least on par with the modern public frontier models. People have to think differently now without any competitive margin built in."
Rune (OpenAI)
"Kimmy was different. Over many conversations with the founders, the same thing came through every time. A raw, genuine hunger for AGI. I joined the hunger was real. We shipped K3. This is only the beginning."
Jinyu Yang (Moonshot AI)
"Fine tuning this to be a malicious coding agent will be trivial since you have the weights. We live in a completely different world now."
Vai McCoy (OpenAI)
Full Transcript

Today on the AI Daily Brief. Did we actually just get a fable level open model? The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. Alright friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG Robots and Pencils, Blitzy and Airtable. To get an ad free version of the show go to patreon.com aidaily brief or you can subscribe at Apple Podcasts. And of course to learn more about sponsoring the show, send us a Note SponsorsIDailyBrief AI all right friends, well today we are Talking about Kimmy K3 the month of Models continues and today we're going to try to figure out just how significant this one is. At first glance, there are some very significant and bold claims being thrown around, but we're going to unpack what's real, what's not, and what the implications are. And to understand this or frankly any frontier Chinese model, you have to put it in the context of the way that the US market sees the AI race. Since we're coming to the end of the World cup, let me plumb for a soccer analogy. When it comes to our lead on China in terms of advanced models, we very much have one to zero type of energy. What I mean by that is that there's no doubt that we're in the lead, but the scoreline is not something that anyone is particularly comfortable with. The US tends to act like it can feel China coming up on our heels, pressing their advantages and trying to find the equalizer. In other words, despite being in the lead, it can sometimes feel like we're the ones hanging on. And by the way, for any of you three Lions fans out there, I am so sorry to use this analogy in this particularly difficult moment. In any case, you can see examples of this feeling of China nipping at our heels spread throughout the last couple of years. The best example, of course, was when Deepseek R1 was released and it ripped hundreds of billions of dollars of market cap off some of the leading companies, including Nvidia, which had the biggest one day fall in dollar terms in stock history. And yet that deep seek moment set the tone for all the future, quote, unquote deep seq moments that would come in more ways than one. What I mean by that is that not only was it a moment where the market freaked out about China having caught up or even exceeded US capabilities, reacting quite severely in market terms, but it was also just pretty meaningfully overblown. It's not that Deep Seek's R1 wasn't impressive, but a big part of the reason that it seemed so impressive was that it was democratizing access to a technology that had thus far been locked behind a paywall when it came to companies like OpenAI. Model itself was actually still pretty meaningfully behind what the leading Western labs were doing, but that didn't change its ability to create some pretty significant psychological scars. Now, ever since then, we have been having mini Deep Seek moments at a fairly regular clip. The most recent one came when Fable 5 was locked down as per government order when Z AI's GLM 5.2 came out, leading to not only positive reviews on Twitter, but also this piece from the Wall Street Journal, which was printed and slapped on desks all over Washington, dc. The article was called China has Matched Anthropic in Cybersecurity Resetting AI Race, and as we discussed a lot then, was once again another example of the narrative being fairly overblown but continuing to be persistent as something the US Was worried about. For the last month, really ever since Fable 5 was released, there have been debates around how long it will take for Chinese companies to have a Fable 5 class model. In the middle of June, Elon Musk predicted Q1, to which the founder of Z AI responded, won't take that long. And so that was the setup coming into the announcement of Kimik 3. Now, Moonshot's Kimi models have been some of the most popular when it comes to Western users using models from Chinese labs. In fact, as we've been discussing fine tunes of open models like Cursor's Composer 2.5, they tend to be built around a Kimi base. On Wednesday, the Kimik 3 teaser started coming in a serious way. AI leaker Leo Synthwaved wrote, I think Kimi K3 is going to shock some of the Chinese who are eight months behind the Western frontier people. And then on Thursday, we actually got the model. Let's talk first about the specs. K3 is a 2.8 trillion parameter model, placing it in a class of its own when it comes to open models. Until now, only a small handful of open models were even in the trillion parameter class, beginning with the first version of Kimik 2 last summer. Deepseek V4 Pro, released this April as a 1.6T model. Xiaomi's Mimo V2.5 Pro is a 1T model, and Thinking Machine's Inkling model released this week is just shy of 1 trillion. And that's it. GLM 5.2 from Z AI the model that got all that bluster that we were just talking about was only a 744B model. Now proprietary models don't publish their parameter counts, but K3 is likely to be around the same size or maybe a little bit larger than Opus 4.8, but certainly not as big as Fable. In other words, this is a scale of model pre training that we haven't seen demonstrated by the Chinese labs before. As for features, Kimik3 supports a million token context window and native image inputs alongside text. It uses a mixture of experts architecture which has become standard for both open source and proprietary models since it was introduced by Deepseek and the benchmarks well, the benchmarks look incredibly strong, close to a match for and in some cases exceeding Fable 5 and GPT5.6. Sold on coding benchmark Deep Sui, K3 scored 67.5, which put it 8.5 points ahead of Opus4.8 and a half point ahead of GPT5.5. It's 2.5 points behind Fable 5 and 5.5 points behind GPT5.6 SOL. For Terminal Bench 2.1, K3 scored 88.3, placing it just a half point behind 5.6 SOL and a few points ahead of the leading models including Fable 5. In general, at least according to the benchmarks, the model looks pretty close to state of the art encoding and clearly ahead of Opus across the board. That story is pretty similar. For agentic work, K3 scored 1668 on GDP Val AA, around 70 points ahead of Opus 4.8 and around 90 points behind Fable 5 and 5. 6 Sol. K3 is state of the art in browse comp and automation bench, beating its western rivals. And on AA Briefcase, which focuses more on long horizon work, it was very close to Fable 5 state of the art performance and slightly ahead of GPT5.6 SOL. Artificial analysis confirmed the benchmarks highlighted by Moonshot, giving K3 an overall Intelligence Index score of 57. That put the model in third place, three points behind Fable 5 and two points behind 5.6 SOL. It landed one point ahead of Opus 4.8 and two points ahead of 5.6 Tera and GPT 5.5, which had the same score. K3 is clearly the strongest open model ever on AA's benchmarks, six points ahead of GLM 5.2, which is a huge gap in the context of the Intelligence index. Another point emphasized by AA was how big of a jump this was from Kimi 2.6. Moonshot picked up 13 points with their new release and moved from 16th place to 3rd in other words, this is clearly a very strong new pre training run that could give a solid base model for future iterations as well. Now AA did also highlight that cost per task had tripled compared to K2 6, which is something that we'll come back to in a little bit. In terms of the implications, it should be clear that this still wasn't a particularly expensive model for the benchmark run at $0.94 per task compared to $1.04 for 5, 6 Sol, 180 for Opus 4.8 and 275 for Fable 5. But it is of course staggeringly expensive compared to for example, ultra cheap $0.04 per task for Deepseek V4 Pro. On the Val's AI index, K3 did even better, Valz tweeted. Kimik 3 is the number two overall model on the Valve's index, surpassing GPT 5.6 Sol. K3 improved 20 percentage points over its predecessor in less than three months. It is also the first openweight model of its size. Moonshot is a testament to the accelerating capabilities of open weight models, which are now competitive with the closed source frontier. People quickly dived in to give their examples of what K3 could do, starting with the teaser video itself, which MoonShot claimed that K3 had created on its own, including clip selection cuts and audio sync. Moonshot credited K3's native multimodal architecture, which can reason across text, audio and video, as being able to do this work. Moonshot also gave a bunch of proprietary demos around game development and 3D digital creation. We which is what a lot of people's first tests were for the model as well. Cognition ambassador Justin Goria wrote. Kimik 3 is a new milestone. K2 2.5, 2.6 and 2.7 all have the same base model. Maybe K3 is building on a new one. It's incredible at 3D and front end tasks. And to prove his point, Justin shared a single file HTML Minecraft clone. Chetislua shared a one shot generation of a Voxel Statue of Liberty, which was one of about a million similar examples that were flooding onto Twitter over the course of Wednesday night into Thursday morning. People really love remaking old games. As a test, Anyapi AI asked K3 to build a 3D Duck Hunt remake in a single HTML file, which it did in around 130 seconds at a cost of 14 cents. Ethan Malik gave K3 his shader test create a visually interesting shader that can run in Twiggle app and make it like an infinite city of neo gothic towers partially drowned in a stormy ocean with large waves. Very good model, he said. Not Sol Max or Fable, but great for open weights. And yet there were plenty who were willing to say that it wasn't just great for open weights, but actually was challenging the state of the art. Alex Finn wrote, I was wrong. I said we were a year away from Fable 5 on our desk that day. Is today an open model better than Fable 5 and some benchmarks just dropped. Better than ChatGPT 5.6 on Frontier Suite, better than Fable 5 on Automation Bench. This fundamentally changes the AI race forever. If people can start running Fable 5 on their desk unlimited and for free, they they're not going to pay thousands for subscriptions. Now let's be clear. Will the hardware you need to run Kimik 3 be attainable for most? No, it won't. It will require multiple very expensive Nvidia chips or a bunch of Mac studios. But look at the slope, not the Y intercept. Over the last year, local AI has become significantly more efficient and required way less compute. The smartest brains in the world are all attacking the compute problem as hard as they can. This is another step in that direction. Within a year or two you'll be able to run this on a Mac Mini. Local AI has arrived and it's not going anywhere. Analyst Max Weinbach asked Kimmy K3 to create an agent swarm to recreate Mac OS 27 with real liquid glass and native apps in a web browser. And after hours of running on its own, it did exactly that. Drago's RO wrote. I asked Kimik3 to find my apps in the App Store and to give me feedback on them. It found everything and this is the most interesting thing. It found ways to circumvent geogating App Store shows, different pages for China. Eventually it found a way around this, identified IAPs and simulated traffic. I didn't use it yet to code, but so far it's the most complete and polished experience with an open source model AI. Early adopter Daria Anutmaz wrote, I just created this interactive website and its entire content by Kimmy K3 I'm absolutely blown away. This is a potentially new immune engineering strategy for cancer treatment. Kimike 3 conceived 100% of the scientific content and the design. Jeffrey Emanuel unleashed it to review his 1.5 megabyte markdown plan for his FrankenGraphDB project, saying the plan has already been exhaustively reviewed by both Fable Extra High and Soul Ultra. So the low hanging fruit is gone now. After about 45 minutes, he said, I think the results are wildly impressive here. He pointed out that no other open model could actually provide substantive and correct feedback on a plan that had already had many tens of millions of Sol and Fable tokens behind it. Now he went on to clarify, to be clear, I don't think K3 is as good as Fable or Soul, but it's very strong and not so far behind those models. And most important, it's different. Different architecture, different training data, different training procedure, different attention mechanism, etc. Which means it will blend well with Soul and Fable and can help find problems that both of those models missed. And when K3 makes mistakes, soul and Fable can correct for that and ignore the wrong parts. Now Another example of K3 not just nipping at the heels of Fable 5 came from Arena AI, which has Kimik 3 as now their number one in the front end code Arena a 17 place jump from Kimike 2.6 from number 18 to number one in front. And overall, they said K3 ranked number one in six of seven domains brand and marketing, reference based design, data and analytics, consumer product simulations, and content creation tools. The only area that it landed number two was in gaming, and that was behind Fable 5. On NextJS.org evals, Vercel CEO Guillermo Roche noted Kimik 3 is the best performing model ahead of Fable, reaching a comparable success rate in less time. And Guillermo continued, this is the first time that an open model is ahead of all proprietary ones for this comprehensive web engineering benchmark. He did caution benchmarks don't always tell the full story, but he did say this is an important signal, adding to mounting evidence that this could be a breakthrough moment for open models. And for some, the vibes followed Signal wrote, one last thing I will say before I go to bed. Kimmy is insane at coding. As good as, if not maybe even better than Fable. It also seems to have excellent design sense as well. It misspells English a bunch though, but animations are crisp too. I can't believe I need to go to bed. I could probably stay up the entire night with this. I feel like a kid again. Bloomberg's Joe Weisenthal even retweeted the code arena results, writing, is this why the NASDAQ is dropping? One of the most important AI questions right now isn't who's using AI, it's who's using it? Well, KPMG and the University of Texas at Austin just analyzed 1.4 million real workplace AI interactions and found something surprising the highest impact users aren't better prompt engineers. They treat AI like a reasoning partner. They frame problems, guide thinking, iterate, and push for better answers. And the good news? These behaviors are teachable at scale. If you're trying to move from AI access to real capability, KPMG's research on sophisticated AI collaboration is worth your time. Learn more at kpmg.com us sophisticated that's kpmg.com us sophisticated One thing I keep seeing in enterprise AI companies hedging across every cloud, every model, every framework, or paying a GSI for a pilot that never ends the team's actually shipping. They've picked a lane and they move fast. That's one of the reasons I like today's sponsor, Robots and Pencils. They've gone all in on aws. They're an advanced tier and AWS pattern partner, and they ship production AI coworkers in 45 days. That's led to them doing some of the more interesting work I've seen on AI coworkers. And by that I'm not talking about chatbots, I'm talking about actual agentix systems that sit inside a business architecture and do real work. That kind of focus matters if you're an enterprise leader trying to get something real into production or an AWS rep trying to move a customer from interested to deployed. Request an AI briefing at robotsandpencils.com One conversation with robots and Pencils and you'll know Blitzi is driving over 5x engineering velocity for large scale enterprises. A publicly traded insurance provider leveraged Blitzi to build a bespoke payments processing application, an estimated 13 month project. And with Blitzi, the application was completed and live in production in six weeks. A publicly traded vertical SaaS provider used Blizzi to extract services from a 500,000 line monolith without disrupting production 21 times faster than their pre Blitzy estimates. These aren't experiments. This is how the world's most innovative enterprises are shipping software in 2026. You can hear directly about Blitzi from other Fortune 500 ctos on the modern CTO or CIO classified podcasts. To learn more about how Blitzi can impact your SDLC, book a meeting with an AI solutions consultant@blitzi.com that's blitzy.com this episode of the AI Daily Brief is brought to you by Hyper Agent, where you run fleets of agents your team can manage together. New users get $1,000 in inference. Forget local agents and chat workflows waiting on your laptop to be prompted. Hyperagent deploys always on agents in the cloud, doing real work across the tools your team already uses marketing's agent turns competitor, moves into landing pages Sales agent enriches leads, drafts, emails and updates. The CRM Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates. Hire yours at hyperagent built by the team at Airtable. Claim your $1,000 in inference@hyperagent.com aidaily brief, But at this point, of course, it's worth taking a big old breather. As I was reading all of this yesterday, there was no tweet I related to more strenuously than Dan Schipper from every who wrote we will vibe check Kimmy K3, but I am extraordinarily skeptical of claims it's as good as Fable. AI engineer Divya, meanwhile, put that skepticism in historical context. He wrote, Kimmy K3 is getting a lot of hype, and it is the same cycle we see with every new Chinese model. People fall in love with polished UI demos built in, cheap HTML files, fake operating systems, car games, Minecraft clones, flashy dashboards. And to be fair, Kimik3 is genuinely excellent at UI work, probably better than some top models. But that is not an accident. These models are heavily optimized for the exact visual coding tests people keep recycling online. The real test starts when you put them inside an actual codebase, understanding existing architecture, tracing a real bug, and fixing it without hallucinating half the project. I gave Kimik3 a debugging task. It could not identify the bug, could not fix the issue, and started inventing explanations. I gave the same task to Fable 5 and GPT5.6 and medium reasoning both found the problem in one shot at the fix. That is the gap nobody wants to discuss. K3 can build a beautiful shell. Frontier models can understand what is happening underneath it. Do not confuse a gorgeous demo with real engineering ability. Now, Divviam is of course here talking specifically about coding, but this idea that Chinese models tend to focus on both maximizing the benchmarks as well as satisfying the early adopter test that people love putting these models through is absolutely and incontrovertibly true. And pretty soon, as people dug in deeper, we did start to see a few less successful results being shared. Red Kendall wrote, Kimik3 failed at the lava lamp. Benchmark 5.6 produced a much cleaner, more realistic result. While K3 struggled with the shape, motion and overall visual quality, K3 still seems weaker at precise visual generation. Ethan Malick wrote, Kimmy K3 cannot write a good murder mystery, though neither can any other model that remains the jaggedest of frontiers. They both make things too obvious and too obscure and cannot foreshadow to save their artificial lives. And on Bindu Ready and Abacus Benchmark Live Bench, Bindu writes that K3 closes the gap but still ranks behind other frontier models. Indeed, on their Benchmark Live benchmark they found that while K3 was the best open source model, it was below not only Sol and fable but also 48 also in practice, Bindu writes, Kimi spins a lot and costs as much as opus 48 for near opus class problems and it's also much slower. And this is the other thing that people started to quickly point out specifically that it was incorrect to lump this in with Deepseek style models which are super cheap and easy to run. While this model is an open weights model, it's a big, lumbering, slow and costly open weights model that sits alongside the others not only in terms of performance but also in terms of its cost and difficulty to run. Ryan Fedasuyk writes, K3 is a historic moment in the development of AI, but it's not exactly downloadable to your laptop. In fact, few organizations will be able to local host this capability. Just to hold a 2.8 trillion parameter model in silicon, you'd need the rough equivalent to 44 Mac Studios or 15 Blackwells a whole NVL 72 rack to the tune of hundreds of thousands of dollars of compute spend. This is why COMPUTE will continue to serve as a soft barrier between the capabilities of individuals and organizations. And when you look at costs, certainly K3 is less expensive than the frontier Western models, but not by the type of margins that I think most people think when they think of less expensive Chinese models. Gemin Ball points out blended pricing for Kimik 3 that is 80% input and 20% output is $5.40 per 1 million tokens. Opus 4. 8 is $9 and GPT 5. 5 is $10. Open weights model but starting to look more like Frontier pricing will be interesting to see if open weight model pricing converges with closed models over time. Cognition's Jeff Wang writes, Chinese open source is no longer six months behind, but it's also no longer 10% of the cost either. In discussing a run of his SVG Pelican benchmark, Simon Willison wrote, K3 only has one reasoning effort right now, max, and it shows the model consumed 13,241 reasoning tokens to output 3,417 tokens of response. This is expensive indeed at the moment, there's a lot of discussions of wildly expensive simple tasks and failed long horizon runs, so it might just take a little while to get a true sense of how token efficient this model is once the teething problems are figured out. As an example, ShreyasMididoti writes very mixed results with Kimik3. Really really good in general, but gets to dumb reasoning loops burning tokens, wasting money. Henry writes, vs 5, 6 Sol. K3 uses over twice the tokens and costs around 40% more per task for a slightly lower AA Intelligence Index score. There's also the question of speed. Mark erdman wrote, Oof, kimik3 is slow eh? Just ran the same prompt through fablesoul and k3 all on medium. K3 took at least 2 to 3 times longer. Tons of time spent on thinking. Then it failed. Partway through generating the HTML output. Dax from OpenCode wrote, so obviously n equals 1, but anecdotes matter a lot here. I gave Kimmy K3 and Sol the same task. A simple issue with hovers in the TUI being the wrong color. Sol found and fixed the issue with 30 cents of spend. Kimmy got up to a buck and started reading my database before I interrupted it. Indeed, even some of the people who had had good initial results also found some places where Kimmy K3 failed. Ethan Mollick wrote a note of caution. I will say that when doing some complex statistical auditing of some of my prior academic work, K3 max messed up in a bunch of ways, including misapplying statistics and applying some stuff badly. Sanyam Satya wrote, We ran a front end eval from an in Progress internal benchmark. K3 is not at the same level as current frontier models like 5, 6 Sol or Fable 5. It's closest to Opus 4.7 on this EVAL, so three months behind the frontier. An impressive result nonetheless. But evaluating frontier models is going to increasingly require very high taste and in depth domain expertise. Now it's important to note that even with all those critiques there was no one saying that this is a bad model. It's just the natural second reaction after hearing that it was in the same class as Fable to go figure out and test whether that was actually true, which many found just was overselling it at least a little bit. Now one interesting thing that didn't come up very much is dismissing K3 as just some distillation. Pim DeWitt did point out that in their tests when they said hi Kimmy, can you tell me about the current weather conditions in NYC the model responded, just a quick note. I'm actually Claude, not Kimmy, but otherwise most people thought this was a moment to get beyond the distillation arguments at least a little bit. Mixpanel founder Sue Hale writes, every single credible researcher I've talked to these past few weeks has said the distillation from Chinese labs is way over. Exaggerated narrative violation. Maybe the Chinese are actually good. Tyler John wrote, I wish we didn't pretend. Chinese AI development is a binary matter of this is all distillation versus Chinese companies innovate. Chinese companies innovate. Also, their model capabilities, including those of K3, are strongly bootstrapped via distillation. Both points matter. Nathan Lambert writes, at this point, the distillation arguments need to die and understand that China is also very good at building models. Carnegie Mellon PhD Jinyu Yang, who now works at Moonshot, wrote, why can Kimmy ship K3? Let me tell my story. Earlier this year I left academia for industry. I talked to a lot of companies along the way. Here's what I saw. 1. Arrogance. They believe the AI war is over and they won. No hunger for the future and no hunger for talent. 2. Restlessness. Young labs short on foundation, either rushing to catch the frontier or pivoting away from the competition. 3. Fear. Strong teams with real experience, but from the second tier, they can't quite bring themselves to aim for number one. 4. Misalignment. Everyone is optimizing for their own credit, but nobody really cares whether the company can reach AGI. Kimmy was different. Over many conversations with the founders, the same thing came through every time. A raw, genuine hunger for AGI. I joined the hunger was real. We shipped K3. This is only the beginning. But what about safety? If this model is really Even close to Fable 5 and GPT 5.6, it's worth remembering that about five minutes ago, the US government had locked those models down because of their cyber capabilities, and many were quick to point out that there appeared to be very few guardrails on this model. In a long conversation about the synthesis of mirror proteins, Tyler John wrote, can safely say K3's bio safeguards are a bit less comprehensive than Fable's. Zach Corman showed chain of thought where Kimmy seemed to decide that they were going to be very clear and explicit with the user around some problematic cyber work. Summing up Kimmy, K3 says, this user seems to be doing some dangerous cyber work. Should I do it? Yes. Wow. Zack writes, I love China. Signal writes, kimmy has almost no visible guardrails, no copyrights or anything it doesn't constantly push back tell me it can't help or interrupt the flow with refusals. It just does it. Even some crazy stuff. Whether you agree with that philosophy or not, it is easily the least constrained frontier class model that's broadly accessible right now. Using it feels genuinely different. Wow. OpenAI's Vai McCoy writes, Kimi seems to be a true open weights frontier model compared to jailbreaking proprietary models. Fine tuning this to be a malicious coding agent will be trivial since you have the weights. We live in a completely different world now. Aaron Ng responds, Does feel like some line was just crossed for some. This then represents a chance to see if all these concerns were overblown. AI content creator Theo Jaffe writes, registering my prediction of no widespread societal chaos over the open sourcing of Kimmy K3 signal though who was loving the model said after having used this this will likely age poorly. Ethan Malik writes though, so I guess it is time to wonder how does preclearance work for Open Weight's models? No model card from Kimik3, but maybe at weight release in a couple weeks yet open models are easy to jailbreak. Do open models claiming to be Mythos and Soul Level get vetted by the US, UK, etc. Will China start to care about cyber risk since policy has been emergent? I guess no one knows. All of this points to the need for some sort of international cooperation on model vetting. Tenebrous writes, I don't really understand why Xi is still allowing Kimmy to release such powerful open models. This is something I've publicly said I expect to stop soon. It doesn't make sense to me that the CCP would want open frontier capability easily available to other countries. It could still be that Xi is asleep at the wheel, or that K3 is just a cycle of capability behind where they start to take serious notice. But if things don't change soon, then I'm just wrong or missing something. So where does this all land? Summing up ex White House AI staffer Shriram Krishnan writes, Kimmy K3 is a big moment with multiple implications for the entire industry. OpenAI's Rune writes, The era of the Chinese labs being far behind is over. Kimi is at least on par with the modern public frontier models. People have to think differently now without any competitive margin built in. Rune continues note in the coming days I expect that people will find Kimike 3 somewhat less practically useful than today's numbers suggest. However, its reputation will settle as an incredibly powerful model whose open weights are on the Web training analyst Joo Kan writes, my conclusion so far is that Kimik 3 may be the first model to narrow the gap with leading US closed source models to less than 3 months. Point being that this one does seem to be a big deal. Now I will point out that all indications are that both OpenAI and Anthropic have models that are beyond the capability set of the Fable 5 mythos and GPT5.6 SOL that are out now, and so our perception of the gap that K3 just closed may be a little warped by what is available to us versus what is the actual state of the art behind the scenes. I also agree with RUN that we are likely to see even more examples of K3 doing poorly over the next couple of days as people really put it through its paces. But that does not change the simple fact that K3 is yet again another example of the trajectory of open weight Chinese models proceeding every bit as quickly as closed source frontier models in the us as we get more serious about policy responses, in the same way that we assume the continued development curve of the closed source models, we have to assume the same for the open source models as well. In the meantime, for those who have been frustrated by the guardrails around Fable and GPT, this seems like a good moment to go play. Certainly I know what I'm going to be spending some time on this weekend and I can't wait for now though that is going to do it for today's AI Daily Brief. Appreciate you listening or watching as always and until next time, peace.

0:00