Was Kimi K3 Built by Copying Fable? The 15-Day Mystery
Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight model topping the coding leaderboard. But with only 15 days between Claude Fable 5 and Kimi K3, accusations of distillation and copying have sparked intense debate.

I was checking Arena scores the other day when Kimi K3 sat at the top of the fullstack coding leaderboard ahead of GPT 5.6 Sol and Claude Fable 5. That caught my attention because Moonshot had just released a 2.8-trillion-parameter open weight model, and the numbers looked promising. Then the accusations started flying.
Within a week of the release, people were saying the model was basically distilled from Anthropic’s Fable. Some were calling it straight theft. I decided to dig into the timeline myself instead of just scrolling past the hot takes. Kimi K3 dropped on July 16. Anthropic’s Claude Fable 5 had become fully available again on July 1 after a short suspension earlier in June. That left about fifteen days between the two. White House science adviser Michael Kratsios posted on July 22 that the US government had information Moonshot distilled Fable to build K3. He said the company built an internal platform that switched between different access methods to stay under the radar.

Anthropic’s public policy head Sarah Heck, backed the claim and called it industrial espionage. Treasury Secretary Scott Bessent discussed possible sanctions and Entity List actions if the pattern continued. None of them released the underlying logs or technical evidence to the public.
I kept coming back to the calendar. Fifteen days is not a lot of time to pull large volumes of outputs, clean them, run supervised fine-tuning, and then do the kind of reinforcement learning that actually moves a frontier model. Braden Hancock from Snorkel AI said the same thing in TechCrunch: you do not get a model this strong that quickly from pure distillation. Nathan Lambert made a similar point if simple output copying was enough, more labs would already be closing the gap the same way and they are not. The full piece is
| Feature | $0 / month | $15 / month | $31 / month | $79 / month | $159 / month |
|---|---|---|---|---|---|
| Annual Billing (Effective Monthly) | $0 / month | $15 / month | $31 / month | $79 / month | $159 / month |
| Agent Concurrent Tasks | 1 task | 2 tasks | 2 tasks | 4 tasks | 4 tasks |
| Agent Priority Queue | × | 4× speed | 4× speed | 4× speed | 4× speed |
| K3 Extra Long Chat Capacity (Up to 1M Tokens) | × | × | × | ✓ | ✓ |
| Customizable Dashboard | ✓ | ✓ | ✓ | ✓ | ✓ |
| Scheduled Tasks | 2 tasks | 10 tasks | 15 tasks | 20 tasks | 25 tasks |
| Widget Task | 2 tasks | 10 tasks | 15 tasks | 20 tasks | 25 tasks |
| Plugins | 15+ types supported | 15+ types supported | 15+ types supported | 15+ types supported | 15+ types supported |
| Swarm | × | ✓ | ✓ | ✓ | ✓ |
| Swarm Running Subtasks | × | 2 subagents | 4 subagents | 8 subagents | 8 subagents |
| Dream Memory | × | ✓ | ✓ | ✓ | ✓ |
| Self-evolving Skill | × | ✓ | ✓ | ✓ | ✓ |
| Goal | × | × | ✓ | ✓ | ✓ |
| Kimi Claw (Web, Android, PC) | × | × | ✓ | ✓ | ✓ |
| Group Chat with Claw | × | × | 10 group chats | 10 group chats | 10 group chats |
| Deploy a Website with a Database | × | ✓ | ✓ | ✓ | ✓ |
Moonshot has not given a detailed training breakdown. One of their employees pointed at the short window and joked that training a frontier model in fifteen days would be Guinness-record territory. The company also released the weights a few days later, which is the opposite of how most “stolen” models behave.
A few screenshots were floating around where Kimi K3 introduced itself as Claude in one conversation. That kind of identity slip happens with distilled models sometimes, but it is not proof by itself. Models pick up phrasing and style from whatever data they see. I have seen smaller open models do the same thing after being fine-tuned on Claude outputs.
Distillation itself is not rare. Almost every lab uses some version of it. The fight is over scale, secrecy, and whether the access violated terms of service. Anthropic had already accused Moonshot, DeepSeek, and MiniMax earlier in the year of running large numbers of accounts against Claude. Those earlier claims involved millions of exchanges. The new accusation ties that pattern specifically to Fable and the K3 release.
Dario Amodei later clarified Anthropic’s position. He said the company has never pushed for a blanket ban on open-weight models. He called non-dangerous open weights a public good. What he wants is tighter control on industrial-scale distillation, continued limits on high-end chips going to China, and mandatory safety testing for capable models whether they are open or closed. That statement came after people accused Anthropic of trying to protect its closed business by shutting down competition.
I keep thinking about what this looks like from the developer side. When an open model shows up that can hold its own on real coding tasks, people download the weights and start using them. The politics around how it was trained matter less once the model is running on your own hardware. At the same time, if the training really did rely on systematic scraping of a closed model’s outputs, that sets a bad precedent. Every lab starts hardening their APIs and the cost of honest iteration goes up.
The evidence so far is mostly official statements and timeline analysis. No public training logs. No side-by-side comparison of internal representations. No clear watermark proof that has been shared. Researchers who looked at the numbers keep saying the fifteen-day window makes a pure Fable-to-K3 distillation story hard to believe. Earlier Claude models could have influenced post-training, but that is a weaker claim than “they copied Fable.”
Moonshot says the gains come from their own architecture work — Kimi Delta Attention and Attention Residuals. Independent checks on the model’s behavior show some original patterns, but also enough surface similarity that the distillation question will not die quietly. For now the mystery sits there. A strong open model appeared, the US government pointed at Anthropic’s latest release as the source, and the calendar makes the simplest version of that story difficult. I will keep watching the Arena scores and any new technical reports. If harder evidence shows up, the conversation will change. Until then, the fifteen days remain the part that does not quite add up.
Share this publication
Related Publications

OpenAI's GPT-5.6 Is Now Optimizing Itself Cutting Costs by 20%
OpenAI released an engineering update on July 29 showing GPT-5.6 Sol optimizing its own serving stack, cutting end-to-end costs by 20% and improving speculative decoding efficiency.

Cursor AI Hits iPad – Mobile Coding Just Got Serious
Cursor launched a dedicated iPad update featuring a tablet-optimized layout, side-by-side agent chats, Apple Pencil markup, PR reviews, and an Inbox status board.