Kimi K3 Weights Are Live and It's the LARGEST Open Model Ever!

Categories: AI

Summary

Moonshot AI released Kimi K3, a 2.88 trillion parameter open-weight model achieving 2.5x intelligence per unit of compute through architectural innovation—not just scale. With 896 experts (16 active per token) and MIT-inspired licensing, it's the largest open model available, but commercial use requires deals above $20M annual revenue.

Key Takeaways

  1. Kimi K3 uses 896 experts with 16 active per token, more than doubling previous expert count. This narrow specialization delivers task-specific expertise instead of generalist models, driving measurable intelligence gains without bloat.
  2. Mixture of Experts architecture innovation achieved 2.5x intelligence-per-compute improvement. Builders should prioritize architectural efficiency over raw parameter scaling for competitive advantage.
  3. MIT-inspired license requires commercial agreements for companies with $20M+ annual revenue or 100M+ users. Startups planning deployment must budget for licensing deals before hitting these thresholds.
  4. Moonshot open-sourced infrastructure rarely released: high-performance attention kernels, MoE communication libraries, and sandbox environments for agent training. Access to training infrastructure is a competitive multiplier for builders.
  5. Kimi K3 benchmarks match Llama 5 Max and GPT-4o, outperforming all other open-weight models. China's open-source approach contrasts Western labs' closed models—expect more capable open alternatives from DeepSeek and GLM teams.

Related topics

Transcript Excerpt

So today the Moonshot team has just released the model weights and technical reports of Kimik K3 which basically means you have access to this model now. You can run it on your own infrastructure and do whatever you want with it. Obviously there are some restrictions but I'll talk about that in a bit. But Kimmy K3 is their most capable model. That's how they're positioning it. It's actually a 2.8 8 trillion mixture of experts model with native visual understanding and a 1 million token context window. They're also positioning this model as the first open 3 trillion class model which tells you that this is a big deal for a lot of us because this is the first time where we're seeing an openweight model make this much stride in a short amount of time. We saw Moonshot AI basically excel their …

More from In The World of AI