VFF - The signal in the noise
News

Nemotron 3 Ultra Matches Closed Models at 10x Lower Cost

Read original
Share
Nemotron 3 Ultra Matches Closed Models at 10x Lower Cost

NVIDIA's Nemotron 3 Ultra model, tuned through LangChain's Deep Agents harness, achieved benchmark-leading performance on agentic AI tasks at one-tenth the inference cost of leading closed models. The optimization came through engineering the orchestration layer rather than retraining the model itself. Companies including Abridge, Amdocs, Box, and EY are already embedding specialized agents built on this stack into their platforms.

  • Nemotron 3 Ultra achieved highest accuracy among open models on LangChain's Deep Agents benchmark while running at 10x lower inference cost than leading closed models
  • Performance gains came from tuning system prompts, tool descriptions, and middleware around the model, not from retraining
  • NVIDIA NemoClaw blueprint packages the tuned stack with LangChain Deep Agents code and NVIDIA OpenShell secure runtime for enterprise deployment
  • Tuned harness is available now through LangChain and hosted on Baseten, Crusoe Cloud, DeepInfra, Fireworks, Nebius, and Together AI

This demonstrates that open-source models can match closed-model performance on complex agentic tasks through better system engineering rather than model scale. The result shifts the economics of enterprise AI by reducing inference costs while maintaining capability, and gives organizations control over their full AI stack from model through runtime.

Enterprises can now run continuous evaluations, experiment faster, and deploy specialized agents at one-tenth the per-run cost of proprietary alternatives. The fully open stack means teams retain ownership and can customize agents for their specific workflows without vendor lock-in.

  • Open models tuned for specific orchestration platforms can achieve parity with closed models on complex tasks, challenging the assumption that proprietary models are necessary for agent performance
  • The economics of agentic AI shift significantly when inference costs drop by 10x, enabling more frequent experimentation and broader deployment across business processes
  • Enterprises increasingly expect ownership and customization of their AI stacks, particularly as agents move from assistive to action-taking roles in core systems

Monitor adoption rates among the named early customers (Abridge, Amdocs, Box, EY) to assess real-world performance and cost savings. Track whether other orchestration platforms follow LangChain's approach of tuning for specific open models, and watch for competitive responses from closed-model providers on pricing and customization.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Telecom Operators Bet on Open AI Models for Network Control
News

Telecom Operators Bet on Open AI Models for Network Control

Telecom operators are adopting open-source AI models as a core strategic component, with 89% of respondents in NVIDIA's State of AI in Telecommunications report citing their importance. Open models enable telcos to customize AI for network operations, customer service, and edge deployment while maintaining control and reducing costs compared to proprietary alternatives. NVIDIA and partners are releasing telecom-specific models like Nemotron 3 Large Telco Model to accelerate this shift.

by Kanika Atri· NVIDIA Blog (AI)
Mistral AI Releases Mistral Large 4 to Challenge Market Leaders
TrendingNews

Mistral AI Releases Mistral Large 4 to Challenge Market Leaders

Mistral AI, a French AI lab, has released Mistral Large 4, a new large multimodal model designed to compete with both American and Chinese AI rivals. The release represents the company's effort to establish itself as a credible alternative in a market dominated by larger players. The model's capabilities and positioning suggest Mistral is targeting both open and closed AI model segments.

by Anna Heim· TechCrunch AI
GLM 5.3 Now Available on Amazon Bedrock
TrendingNews

GLM 5.3 Now Available on Amazon Bedrock

GLM 5.3, a 753-billion-parameter mixture-of-experts model from Zhipu AI, is now available on Amazon Bedrock with managed APIs and cross-region inference. The model is optimized for coding and long-horizon agentic tasks, with reported improvements in coding benchmarks and emergent cybersecurity capabilities. Enterprise customers can access it without managing infrastructure, with support for prompt caching and OpenAI-compatible APIs.

by Alex Thewsey· AWS Machine Learning Blog
OpenAI releases GPT-6 implementation guide for startups
News

OpenAI releases GPT-6 implementation guide for startups

OpenAI has published a practical guide for startups building with GPT-6 models, covering model selection, reasoning effort tuning, prompt optimization, tool coordination, and production workflow preparation. The guide addresses the operational and technical decisions required to deploy GPT-6 effectively in startup environments. It reflects growing focus on helping developers move beyond basic API usage to production-ready implementations.

· OpenAI