VFF - The signal in the noise
News

Meta Muse Voice Transcribe: $0.18/hour real-time diarization for 20+ speakers

Read original
Share
Meta Muse Voice Transcribe: $0.18/hour real-time diarization for 20+ speakers

Meta has launched Muse Voice Transcribe, a real-time speech-to-text model priced at $0.18 per hour of audio that combines transcription, speaker diarization for 20+ speakers, and endpoint detection in a single model. The system supports over 70 languages, with 25 extensively validated for release, and handles multilingual code-switching without separate post-processing. While competitors like Speechmatics support higher speaker counts (up to 100), Meta's combination of low latency, high-capacity diarization, and aggressive pricing targets enterprise developers building meeting systems, call analytics, and ambient AI applications.

  • Meta Muse Voice Transcribe priced at $0.18 per hour of processed audio
  • Supports real-time diarization for 20+ speakers with integrated speaker attribution in token sequence
  • Trained on 70+ languages, 25 extensively validated, handles seamless code-switching
  • Competitors Speechmatics (50-100 speakers) and Amazon Transcribe (30 speakers) offer higher speaker-count ceilings

Diarization, the ability to identify who said what, is becoming critical infrastructure for AI systems processing meeting transcripts, customer service calls, and compliance workflows. Misattributing statements to wrong participants can render corporate records unreliable even when transcription is accurate. Meta's integration of speaker attribution directly into its model architecture rather than as post-processing represents a shift in how voice AI handles multi-speaker scenarios.

For enterprises building meeting assistants, call analytics platforms, and ambient AI systems, Muse's combination of real-time diarization, low latency, and low API pricing ($0.18/hour) reduces infrastructure complexity and cost compared to piecing together separate transcription and diarization services. The model's ability to handle 20+ speakers in real-time addresses a core pain point in multi-participant business communications without requiring external speaker clustering pipelines.

  • Meta is positioning itself as a cost-competitive alternative in the speech-to-text market, undercutting traditional vendors on price while bundling capabilities previously sold separately
  • Integration of diarization as a first-class feature in the model architecture signals industry movement toward treating speaker attribution as core rather than auxiliary functionality
  • Enterprise adoption will likely depend less on raw speaker-count maximums and more on latency, accuracy, and total cost of ownership for real-time multi-speaker scenarios

Monitor whether Meta's aggressive pricing and integrated architecture drive adoption among enterprise developers, and whether competitors respond with price adjustments or architectural changes. Track real-world performance metrics on diarization accuracy and latency in production meeting and call-center environments, as published specifications do not always reflect practical performance under load.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Meta Launches Muse Voice Transcribe Audio AI Model
TrendingModel Release

Meta Launches Muse Voice Transcribe Audio AI Model

Meta unveiled Muse Voice Transcribe, a new audio AI model that transcribes speech to text and segments audio by speaker. CEO Mark Zuckerberg announced the model on Threads, noting it was trained on more than 70 hours of audio data. The model represents Meta's continued push into audio AI capabilities alongside its existing generative AI portfolio.

by Jyoti Mann· The Information
Clipto hits $250M valuation on path to profitability
TrendingNews

Clipto hits $250M valuation on path to profitability

Clipto, a three-year-old AI startup that uses machine learning to search large video datasets, has reached a $250 million valuation after raising $15 million in new funding. The company achieved $15 million in annual recurring revenue and profitability before closing the round. The funding reflects investor confidence in AI-powered video search as a commercial tool for handling terabytes of media.

by Kate Park· TechCrunch AI
Google Launches Gemini 3.5 Transcribe with 85+ Language Support
TrendingModel Release

Google Launches Gemini 3.5 Transcribe with 85+ Language Support

Google has released Gemini 3.5 Transcribe, a new transcription model that automatically detects specialized jargon and supports more than 85 languages. The model represents an improvement over its predecessor, Chirp 3, with better multilingual performance and lower wording error rates. Users can edit transcriptions using voice commands. The release comes as Google continues to roll out updates to its Gemini Audio suite while the promised Gemini 3.5 Pro model remains unreleased since its June launch window.

by Jess Weatherbed· The Verge AI
Particle's Radar makes 130K podcasts searchable for AI

Particle's Radar makes 130K podcasts searchable for AI

Particle has launched a podcast intelligence platform called Radar that transcribes and analyzes over 130,000 podcasts, making their content searchable on the web and accessible to AI agents via API and MCP. The platform enables both human users and AI systems to query podcast conversations at scale. This addresses a significant gap in AI training data and search accessibility for audio content.

by Sarah Perez· TechCrunch AI