VFF - The signal in the noise
News

Atlantic Maps Four Music Datasets Powering AI Models

Read original
Share
Atlantic Maps Four Music Datasets Powering AI Models

The Atlantic's Alex Reisner has created a searchable public database of four music datasets used to train AI models, including two massive collections of 12 million and 9 million tracks. The datasets have been downloaded thousands of times, with Google and Stability AI confirming their use in research papers. The discovery highlights the scale of music data being fed into AI systems and raises questions about artist consent and compensation.

  • The Atlantic identified and made searchable four datasets containing music used to train AI models
  • Two datasets contain 12 million and 9 million tracks respectively, with two others holding over 100,000 songs each
  • Google and Stability AI have confirmed using these datasets in published research
  • The datasets have been downloaded thousands of times, though exact usage remains difficult to track

This disclosure exposes the scale and sources of music data powering generative AI systems, a critical gap in transparency around AI training practices. Artists and rights holders have limited visibility into whether their work is being used to train commercial AI models, making this database a rare window into the actual data fueling the industry.

Companies developing music-generating AI and other generative models rely on large-scale training datasets, often sourced from public or semi-public repositories. Understanding which datasets are in use helps stakeholders track potential licensing and rights issues, while also revealing competitive intelligence about training approaches.

  • The music industry lacks effective mechanisms to track and control use of artist work in AI training, creating ongoing legal and ethical exposure for AI developers
  • Public datasets remain a primary source for AI training despite growing scrutiny, suggesting regulatory frameworks have not yet constrained data sourcing practices
  • Transparency tools like this database may become necessary for artists and rights holders to identify and challenge unauthorized use of their work

Monitor whether this disclosure prompts legal action from artists or rights organizations against companies confirmed to have used these datasets. Watch for industry responses around data licensing standards and whether AI developers shift toward licensed or proprietary training data sources.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Meta Deploys Thousands of Engineers to Train Coding AI
TrendingNews

Meta Deploys Thousands of Engineers to Train Coding AI

Meta is deploying its in-house coding agent MetaCode to thousands of engineers to improve the coding capabilities of its AI models and close the gap with Anthropic and OpenAI. VP Maher Saba has asked engineers to submit at least one code change per week for review and integration. The feedback loop has already improved Meta's latest model, Muse Spark 1.1, and will be used to train an upcoming model called Watermelon.

by Jyoti Mann· The Information
AI Drug Discovery Hits a Data Wall
TrendingNews

AI Drug Discovery Hits a Data Wall

AI is accelerating drug discovery by enabling predictive design of candidates and hit identification at scale, but the technology is exposing critical gaps in data quality and lab infrastructure. Drug companies are hitting a 'data wall' where publicly available datasets lack the structure and diversity needed to train accurate models, while lab teams struggle to validate the growing volume of AI-generated compounds. Success depends on closing the loop between computational prediction and experimental validation through better data collection and integration.

by MIT Technology Review Insights· MIT Technology Review
Brain Waves Join Video as Physical AI Training Data
TrendingNews

Brain Waves Join Video as Physical AI Training Data

Frontier physical AI models are moving beyond video training data to incorporate multiple camera angles, dense annotation, and brain wave readings as training inputs. The shift reflects growing recognition that traditional video datasets alone are insufficient for training AI systems that interact with the physical world. Brain wave data represents an emerging frontier in multimodal training approaches for robotics and embodied AI.

by Tim Fernholz· TechCrunch AI
Mercor's $614M Revenue Surge Hinges on AI Lab Spending

Mercor's $614M Revenue Surge Hinges on AI Lab Spending

Mercor, a three-year-old data startup that trains AI models through contractor networks, generated $614 million in gross revenue in the first half of 2026, up 70% from all of 2025. The company's growth is heavily concentrated among AI foundation model makers, with 91% of first-half revenue coming from customers like OpenAI, Anthropic, and Google DeepMind. This revenue concentration reveals both the startup's market traction and its dependency on a narrow customer base.

by Julia Hornstein· The Information