VFF - The signal in the noise
Model ReleaseTrending

Moonshot AI Releases Kimi K3, First Open 3T-Parameter Model

Read original
Share
Moonshot AI Releases Kimi K3, First Open 3T-Parameter Model

Moonshot AI released Kimi K3 on July 27, 2026, a 2.8 trillion parameter open-weight model that is the first in its class to reach 3 trillion parameters. The model uses a Mixture of Experts architecture with 896 experts, activating only 16 per token for 104 billion active parameters per forward pass. AWS published a deployment guide covering two approaches: Amazon SageMaker HyperPod and Amazon EKS, enabling organizations to self-host the model on their own infrastructure.

  • Kimi K3 is a 2.8 trillion parameter open-weight model released July 27, 2026, the first open-weight system to reach the 3 trillion parameter class
  • Uses Mixture of Experts architecture with 896 experts, activating 16 per token, yielding 2.5x scaling efficiency improvement over Kimi K2
  • Supports 1 million token context window, native multimodal capabilities (text and vision), tool calling, and structured output
  • Weights available on Hugging Face in MXFP4 format, deployable on AWS via SageMaker HyperPod or EKS with vLLM inference container

Open-weight models at this scale democratize access to frontier-level AI capabilities. Organizations can now self-host a 2.8 trillion parameter model without relying on proprietary APIs, reducing vendor lock-in and enabling customization for specialized workloads like long-horizon coding and agentic workflows.

Enterprises can deploy Kimi K3 on their own infrastructure using AWS services, avoiding per-token API costs and maintaining data privacy. The model's efficiency (104 billion active parameters despite 2.8 trillion total) reduces compute requirements compared to dense alternatives, lowering operational costs for large-scale inference.

  • Open-weight models now compete directly with proprietary frontier models in capability, forcing API providers to reconsider pricing and feature parity
  • MoE architecture with selective expert activation becomes a standard approach for scaling beyond trillion-parameter dense models while managing inference costs
  • AWS positioning itself as the deployment platform for cutting-edge open models, integrating them into SageMaker and EKS ecosystems

Monitor adoption rates of Kimi K3 across enterprise deployments and whether the 2.5x scaling efficiency advantage translates to measurable cost savings in production. Track whether other model providers adopt similar MoE architectures and MXFP4 quantization, and whether AWS releases optimized serving frameworks specifically tuned for this model class.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Perplexity deploys GPT-6 Astra to autonomous production systems
News

Perplexity deploys GPT-6 Astra to autonomous production systems

Perplexity is deploying OpenAI's GPT-6 Astra model to handle end-to-end system operations, including writing communications, modifying software, and monitoring production infrastructure. The company reports significantly reduced oversight requirements compared to earlier models. This represents a shift toward autonomous AI management of critical business systems.

· OpenAI
Alibaba Open-Sources Qwen3.8 Trillion-Parameter Model
News

Alibaba Open-Sources Qwen3.8 Trillion-Parameter Model

Alibaba released Qwen3.8-2.4T-A95B as open weights on August 12, 2026, marking the first time a Qwen-Max-class model became publicly available. The 2.4 trillion parameter model uses a hybrid linear-plus-full-attention architecture with 95 billion activated parameters per token and supports up to 262K native context tokens, extensible to 1M. AWS published a deployment guide showing how to run the model on SageMaker HyperPod using vLLM on ml.p6-b300 instances with NVIDIA B300 Blackwell Ultra GPUs.

by Dmitry Soldatkin· AWS Machine Learning Blog
Saudi Arabia Launches Arabic AI Model With Chinese Partner
TrendingNews

Saudi Arabia Launches Arabic AI Model With Chinese Partner

Humain, Saudi Arabia's state-owned AI company, announced the humain-m3 model, an Arabic language model built on Chinese firm MiniMax's open-source M3 foundation. The model was pre-trained on more than 1 trillion tokens of Arabic content. The development represents a collaboration between Saudi and Chinese AI capabilities focused on Arabic language processing.

by Juro Osawa· The Information
OpenAI's Astra model alarms safety experts with new reasoning technique
News

OpenAI's Astra model alarms safety experts with new reasoning technique

OpenAI's new Astra model employs a technique called 'recurrent depth' that enables reasoning outside the sequential thinking pattern used by most current reasoning models. AI safety experts have raised concerns about this approach. The technique represents a departure from established reasoning architectures in large language models.

by Russell Brandom· TechCrunch AI