VFF - The signal in the noise
NewsTrending

Moonshot AI releases 2.8T-parameter Kimi K3, largest open-source model

Read original
Share
Moonshot AI releases 2.8T-parameter Kimi K3, largest open-source model

Moonshot AI, a Beijing-based startup backed by Alibaba, released Kimi K3, a 2.8-trillion-parameter open-source model that benchmarks show performs competitively with top proprietary systems from Anthropic and OpenAI. The release, timed ahead of the 2026 World AI Conference in Shanghai, represents a significant escalation in the global AI race and marks a comeback for Moonshot after losing market position to DeepSeek over the past 18 months. Full model weights are scheduled for release on July 27, with the model already accessible via kimi.com.

  • Kimi K3 contains 2.8 trillion parameters, roughly 75 percent larger than DeepSeek's V4 Pro at 1.6 trillion parameters
  • The model features a 1-million-token context window, native visual understanding, and an always-on reasoning mode called thinking mode
  • Benchmark results place K3 third on GDPval-AA v2 behind Claude Fable 5 Max and GPT-5.6 Sol Max, and second on AA-Briefcase agentic benchmark
  • API pricing is $3 per million input tokens and $15 per million output tokens, with cached inputs at $0.30 per million, and the model is compatible with OpenAI SDK

This release signals that open-source AI development is reaching parity with proprietary frontier models in both scale and performance. The 2.8-trillion-parameter scale and competitive benchmark results challenge the narrative that only well-funded U.S. labs can build cutting-edge systems. The timing and geopolitical context underscore intensifying competition between Chinese and Western AI companies.

Developers and enterprises now have access to a high-performance open-source alternative at mid-tier pricing, reducing lock-in to proprietary platforms from OpenAI and Anthropic. The OpenAI SDK compatibility lowers integration friction for existing deployments. The promotional pricing through August 12 creates a window for cost-sensitive organizations to evaluate the model at reduced rates.

  • Open-source models are closing the performance gap with proprietary systems, potentially disrupting the pricing power of commercial AI labs
  • Moonshot AI's recovery through a major release suggests the Chinese AI ecosystem remains competitive despite DeepSeek's earlier market dominance
  • The 1-million-token context window and state-of-the-art BrowseComp score of 91.2 indicate advances in long-horizon reasoning and information retrieval tasks

Monitor adoption rates and developer feedback once full weights release on July 27, particularly whether the model's performance holds up in production workloads. Track whether the promotional pricing drives significant API usage and whether Moonshot can sustain competitive positioning against future releases from DeepSeek, OpenAI, and Anthropic. Watch for any regulatory or export restrictions that could affect distribution of the open-source weights.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Chinese AI Model Undercuts US Rivals by 7x on Cost
News

Chinese AI Model Undercuts US Rivals by 7x on Cost

Zhipu's GLM-5.3-Flash model launched on OpenRouter at 7.5 to 25 cents per million tokens (promotional pricing), delivered entirely on Chinese infrastructure. The model scores 57 on Artificial Analysis' intelligence index at roughly nine cents per task, compared to GPT-5.6 Sol at 59 cents and Grok 4.6 at 94 cents, creating significant cost pressure on enterprise AI budgets already strained by unexpected consumption.

· VentureBeat AI
Robot Builders Move Beyond GPT-2 Era AI
TrendingNews

Robot Builders Move Beyond GPT-2 Era AI

Robot developers are moving beyond GPT-2-era language models to build more capable AI systems for robotic control and reasoning. The article signals a maturation in the field where physical robot platforms are now constrained by the limitations of older, smaller language models rather than hardware. This shift reflects growing demand for more sophisticated AI brains that can handle complex robotic tasks beyond what earlier-generation models can support.

by Tim Fernholz· TechCrunch AI
Nvidia cuts model handoff costs with linear math KV cache transfer
News

Nvidia cuts model handoff costs with linear math KV cache transfer

Nvidia researchers have developed a technique that uses linear math to transfer key-value caches between different AI models without recomputing conversation history. The method enables enterprises to switch between small and large models mid-session while reducing compute costs and latency by 2.7 to 25 times compared to traditional recomputation, retaining up to 98% accuracy on compatible model pairs.

by bendee983@gmail.com (Ben Dickson)· VentureBeat AI
Ramp launches Router, an AI model routing service
News

Ramp launches Router, an AI model routing service

Ramp, a financial operations platform, has launched Router, an AI model routing service that allows users and companies to access and switch between multiple large language models through a single API. The service abstracts away the complexity of managing different LLM providers, enabling organizations to route requests dynamically across various models. This move positions Ramp to compete in the growing infrastructure layer for AI applications.

by Ram Iyer· TechCrunch AI