VFF - The signal in the noise
NewsTrending

Nvidia Eyes Data Labeling Investment as Open-Source AI Ambitions Grow

Read original
Share
Nvidia Eyes Data Labeling Investment as Open-Source AI Ambitions Grow

Nvidia is in discussions to invest in Mercor, a data labeling company, as part of a $20 billion funding round led by existing investor General Catalyst. Mercor has historically served closed-source AI model makers like OpenAI, Google, and Anthropic, but revenue from Nvidia is growing as the chip designer develops its Nemotron open-source models. The investment signals Nvidia's commitment to competing in open-source AI model development.

  • Nvidia discussing investment in Mercor, a data labeling provider, at $20 billion valuation
  • General Catalyst leading the funding round as existing investor
  • Mercor's revenue from Nvidia growing as chip firm develops Nemotron open-source models
  • Mercor historically relied on closed-source AI makers like OpenAI, Google, and Anthropic for revenue

Nvidia's potential investment in Mercor reveals the infrastructure dependencies underpinning AI model development. Data labeling is a critical bottleneck in training high-quality AI systems, and Nvidia's move to fund a supplier suggests the company is building vertical integration around its open-source AI ambitions. This also indicates Nvidia sees open-source models as strategically important enough to secure dedicated data resources.

For Mercor, Nvidia's investment validates its business model and provides a major customer anchor as it diversifies beyond closed-source model makers. For Nvidia, securing data labeling capacity reduces dependency on third-party suppliers and supports its Nemotron models' competitive positioning against other open-source offerings. The $20 billion valuation reflects investor confidence in data labeling as a core AI infrastructure business.

  • Nvidia is moving beyond chip design into the broader AI model development supply chain, including data infrastructure
  • Open-source AI model development is becoming capital-intensive, requiring dedicated data labeling resources at scale
  • Mercor's customer base is shifting from primarily closed-source model makers toward open-source developers like Nvidia

Monitor whether Nvidia completes the investment and at what final valuation. Track how Mercor's revenue mix evolves between closed-source and open-source customers, and whether other chip makers or model developers follow Nvidia's lead in funding data infrastructure suppliers. Watch for announcements about Nemotron model performance and adoption, which would validate the value of Mercor's data labeling work.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

ChatGPT Now Tracks Your Keystrokes on macOS

ChatGPT Now Tracks Your Keystrokes on macOS

OpenAI has introduced Computer History, a new feature in ChatGPT's macOS desktop app that tracks user clicks and keystrokes to build activity timelines for AI reference. The feature is opt-in and allows users to exclude specific apps and websites, with automatic filtering of incognito and private browsing content. This capability enables ChatGPT to suggest automations and resume incomplete tasks based on observed user behavior.

by Terrence O’Brien· The Verge AI
Startup Slack Threads Become Commodity for AI Training
TrendingNews

Startup Slack Threads Become Commodity for AI Training

AI training companies like Mercor are actively acquiring internal communications and code from startups, offering payments up to $300,000 for Slack threads, GitHub records, and meeting transcripts. Warmly's CEO received four such acquisition offers within days of the company's HubSpot acquisition announcement. The practice highlights how internal startup data has become a commodity for AI model training, even as acquirers may not want the same datasets.

by Alix Coutures· The Information
Data Infrastructure, Not AI Models, Limits Agent Success

Data Infrastructure, Not AI Models, Limits Agent Success

A MIT Technology Review Insights report based on a survey of 300 data and technology executives finds that legacy data systems are a major blocker to AI agent adoption and effectiveness. Organizations with mature data infrastructure, termed 'data leaders,' report significantly higher trust in agent decisions and fewer scaling constraints than 'data laggards.' The research suggests that without modernizing data systems, enterprises will struggle to realize ROI from agentic AI despite widespread adoption plans.

by MIT Technology Review Insights· MIT Technology Review
Meta Deploys Thousands of Engineers to Train Coding AI
TrendingNews

Meta Deploys Thousands of Engineers to Train Coding AI

Meta is deploying its in-house coding agent MetaCode to thousands of engineers to improve the coding capabilities of its AI models and close the gap with Anthropic and OpenAI. VP Maher Saba has asked engineers to submit at least one code change per week for review and integration. The feedback loop has already improved Meta's latest model, Muse Spark 1.1, and will be used to train an upcoming model called Watermelon.

by Jyoti Mann· The Information