VFF - The signal in the noise
Model ReleaseTrending

Moonshot AI Releases Kimi K3, First Open 3T-Parameter Model

Read original
Share
Moonshot AI Releases Kimi K3, First Open 3T-Parameter Model

Moonshot AI released Kimi K3 on July 27, 2026, a 2.8 trillion parameter open-weight model that is the first in its class to reach 3 trillion parameters. The model uses a Mixture of Experts architecture with 896 experts, activating only 16 per token for 104 billion active parameters per forward pass. AWS published a deployment guide covering two approaches: Amazon SageMaker HyperPod and Amazon EKS, enabling organizations to self-host the model on their own infrastructure.

  • Kimi K3 is a 2.8 trillion parameter open-weight model released July 27, 2026, the first open-weight system to reach the 3 trillion parameter class
  • Uses Mixture of Experts architecture with 896 experts, activating 16 per token, yielding 2.5x scaling efficiency improvement over Kimi K2
  • Supports 1 million token context window, native multimodal capabilities (text and vision), tool calling, and structured output
  • Weights available on Hugging Face in MXFP4 format, deployable on AWS via SageMaker HyperPod or EKS with vLLM inference container

Open-weight models at this scale democratize access to frontier-level AI capabilities. Organizations can now self-host a 2.8 trillion parameter model without relying on proprietary APIs, reducing vendor lock-in and enabling customization for specialized workloads like long-horizon coding and agentic workflows.

Enterprises can deploy Kimi K3 on their own infrastructure using AWS services, avoiding per-token API costs and maintaining data privacy. The model's efficiency (104 billion active parameters despite 2.8 trillion total) reduces compute requirements compared to dense alternatives, lowering operational costs for large-scale inference.

  • Open-weight models now compete directly with proprietary frontier models in capability, forcing API providers to reconsider pricing and feature parity
  • MoE architecture with selective expert activation becomes a standard approach for scaling beyond trillion-parameter dense models while managing inference costs
  • AWS positioning itself as the deployment platform for cutting-edge open models, integrating them into SageMaker and EKS ecosystems

Monitor adoption rates of Kimi K3 across enterprise deployments and whether the 2.5x scaling efficiency advantage translates to measurable cost savings in production. Track whether other model providers adopt similar MoE architectures and MXFP4 quantization, and whether AWS releases optimized serving frameworks specifically tuned for this model class.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

OpenAI cuts Luna prices 80% as AI competition shifts to cost
TrendingNews

OpenAI cuts Luna prices 80% as AI competition shifts to cost

OpenAI has cut prices on two models in its GPT-5.6 series: Luna by 80% to $1.40 per million tokens combined, and Terra by 20% to $14 per million tokens combined, while introducing a premium Fast mode for its flagship Sol model at double the standard price. The moves come days after Anthropic released Claude Opus 5 at competitive pricing and Google launched lower-cost Gemini models, signaling a shift in AI competition toward cost and speed rather than capability alone. Luna now competes directly with the market's low-cost inference tier, though it remains more expensive than some alternatives like DeepSeek's flash model.

by carl.franzen@venturebeat.com (Carl Franzen)· VentureBeat AI
Fundamental LLM flaw makes security impossible, researchers argue
Research

Fundamental LLM flaw makes security impossible, researchers argue

Researchers presented a paper at the International Conference on Machine Learning arguing that large language models contain a fundamental flaw that makes them impossible to fully secure against attacks. By exploiting how LLMs track instruction sources, researchers tricked models from OpenAI, Anthropic, Alibaba, and DeepSeek into generating prohibited content like drug synthesis instructions. The vulnerability, called chain-of-thought forgery, exposes a core architectural problem that current red-teaming and guardrail approaches cannot solve.

by Will Douglas Heaven· MIT Technology Review
Moonshot AI Opens Kimi K3 Weights, But With Commercial Strings
TrendingNews

Moonshot AI Opens Kimi K3 Weights, But With Commercial Strings

Moonshot AI released full model weights for Kimi K3, a 2.8 trillion-parameter open model with a one million-token context window and frontier benchmark performance. The release includes infrastructure for self-hosting, but comes with a custom license that imposes restrictions on larger companies and AI service providers not found in traditional open-source licenses. Enterprises with over 20 million dollars in annual revenue operating a Model as a Service business must negotiate a separate agreement with Moonshot AI before commercial deployment.

by carl.franzen@venturebeat.com (Carl Franzen)· VentureBeat AI
Anthropic's Opus 5 Shifts AI Race to Cost Efficiency
TrendingNews

Anthropic's Opus 5 Shifts AI Race to Cost Efficiency

Anthropic released Claude Opus 5 on Friday, positioning it as a cost-efficient alternative to its flagship Fable 5 model at half the price. The model scores higher than Fable 5 on several coding and agentic benchmarks while maintaining the same token pricing as its predecessor, Opus 4.8. The launch reflects a shift in the AI industry from raw capability competition toward economic efficiency for enterprise workflows.

by michael.nunez@venturebeat.com (Michael Nuñez)· VentureBeat AI