Moonshot AI Releases Kimi K3, First Open 3T-Parameter Model

Moonshot AI released Kimi K3 on July 27, 2026, a 2.8 trillion parameter open-weight model that is the first in its class to reach 3 trillion parameters. The model uses a Mixture of Experts architecture with 896 experts, activating only 16 per token for 104 billion active parameters per forward pass. AWS published a deployment guide covering two approaches: Amazon SageMaker HyperPod and Amazon EKS, enabling organizations to self-host the model on their own infrastructure.
TL;DR
- Kimi K3 is a 2.8 trillion parameter open-weight model released July 27, 2026, the first open-weight system to reach the 3 trillion parameter class
- Uses Mixture of Experts architecture with 896 experts, activating 16 per token, yielding 2.5x scaling efficiency improvement over Kimi K2
- Supports 1 million token context window, native multimodal capabilities (text and vision), tool calling, and structured output
- Weights available on Hugging Face in MXFP4 format, deployable on AWS via SageMaker HyperPod or EKS with vLLM inference container
Why It Matters
Open-weight models at this scale democratize access to frontier-level AI capabilities. Organizations can now self-host a 2.8 trillion parameter model without relying on proprietary APIs, reducing vendor lock-in and enabling customization for specialized workloads like long-horizon coding and agentic workflows.
Business Impact
Enterprises can deploy Kimi K3 on their own infrastructure using AWS services, avoiding per-token API costs and maintaining data privacy. The model's efficiency (104 billion active parameters despite 2.8 trillion total) reduces compute requirements compared to dense alternatives, lowering operational costs for large-scale inference.
Key Implications
- Open-weight models now compete directly with proprietary frontier models in capability, forcing API providers to reconsider pricing and feature parity
- MoE architecture with selective expert activation becomes a standard approach for scaling beyond trillion-parameter dense models while managing inference costs
- AWS positioning itself as the deployment platform for cutting-edge open models, integrating them into SageMaker and EKS ecosystems
What to Watch
Monitor adoption rates of Kimi K3 across enterprise deployments and whether the 2.5x scaling efficiency advantage translates to measurable cost savings in production. Track whether other model providers adopt similar MoE architectures and MXFP4 quantization, and whether AWS releases optimized serving frameworks specifically tuned for this model class.
Subscribe to the newsletter
The latest stories and analysis, delivered to your inbox.
Free. No spam. Unsubscribe any time.



