VFF - The signal in the noise
News

Kimi K3 Arrives on AWS Bedrock with 1M Token Context

Read original
Share
Kimi K3 Arrives on AWS Bedrock with 1M Token Context

AWS has made Kimi K3, a 2.8 trillion parameter open-weight model from Moonshot AI, available on Amazon Bedrock. The model features native vision capabilities, a 1-million-token context window, and support for explicit prompt caching. It is positioned for long-running coding and knowledge work tasks that require sustained context across large documents and repositories.

  • Kimi K3 is now available on Amazon Bedrock with 2.8 trillion parameters and native vision capabilities
  • The model offers a 1-million-token context window and approximately 2.5x improvement in scaling efficiency over Kimi K2
  • Kimi K3 is the first open-weight model on Bedrock to support explicit prompt caching, reducing latency and input costs
  • AWS has added dozens of open-weight models to Bedrock since 2025, with platform-level capabilities like tool calling and structured output

Open-weight models are shifting AI economics by allowing organizations to match workloads to the right balance of capability, speed, and cost. Kimi K3's large context window and efficiency gains make it practical for sustained work on complex coding and knowledge tasks. The availability of such models on managed platforms like Bedrock lowers barriers to production deployment.

Companies can now adopt high-capability open models without building custom infrastructure or changing security practices. AWS's data boundary protections, zero data retention, and zero operator access mean teams can use Kimi K3 for sensitive work while maintaining compliance and data control. The model's prompt caching feature directly reduces inference costs for repeated context usage.

  • Open-weight models are becoming viable alternatives to proprietary models for enterprise workloads, particularly in coding and knowledge work
  • Managed platforms like Bedrock are consolidating access to multiple open models with consistent security and API standards, reducing vendor lock-in concerns
  • The 1-million-token context window enables new use cases around long-document analysis and large codebase understanding that were previously impractical

Monitor whether Kimi K3 adoption accelerates on Bedrock and whether other open-weight model providers follow with similarly large context windows. Watch for performance comparisons between Kimi K3 and proprietary models on coding and knowledge tasks. Track whether prompt caching becomes a standard feature across other open models on Bedrock.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Tsinghua-Founded Naive AI Hits $1.4B Valuation Before LLM Launch
News

Tsinghua-Founded Naive AI Hits $1.4B Valuation Before LLM Launch

Naive AI, a Beijing-based startup founded by a Tsinghua University professor in February, has reached a $1.4 billion valuation after raising $400 million across three funding rounds from investors including Tencent. The company plans to release its first large language model this month as an open-weight model, positioning itself as a new entrant in China's competitive LLM market alongside DeepSeek, Moonshot, and Alibaba.

by Juro Osawa· The Information
PrismML Bets on Compact LLMs to Reshape AI Deployment
News

PrismML Bets on Compact LLMs to Reshape AI Deployment

PrismML, an AI lab, is developing a compact large language model intended to shift how AI is deployed and used. The article positions PrismML as an emerging player worth attention in the AI space, though specific technical details, capabilities, or business model are not provided in the source material.

by Julie Bort· TechCrunch AI
Vera Rubin NVL72 Debuts With 3.7x Throughput Gain Over GB300
Research

Vera Rubin NVL72 Debuts With 3.7x Throughput Gain Over GB300

NVIDIA's Vera Rubin NVL72 system achieved leading performance in its MLPerf Inference v6.1 debut, delivering up to 3.7x higher throughput than the GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1. The results demonstrate the effectiveness of full-stack hardware and software codesign, including enhanced Tensor Cores, NVFP4 precision, and disaggregated serving techniques. A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, and software optimizations alone delivered up to 1.6x performance gains from v6.0 to v6.1.

by Zhihan Jiang· NVIDIA Blog (AI)
Lightweight dual-model agents show promise for autonomous materials research
Research

Lightweight dual-model agents show promise for autonomous materials research

Researchers at Nature Machine Intelligence have demonstrated a dual-model architecture for autonomous crystal materials research using two lightweight large language models working collaboratively. The approach combines reasoning and scientific tool execution while maintaining computational efficiency and local deployability. The method achieves competitive performance without requiring expensive infrastructure, making advanced materials research more accessible.

by Tongyu Shi· Nature Machine Intelligence