VFF - The signal in the noise
News

Loka Cuts Voice AI Latency with Amazon Nova 2 Sonic

Read original
Share
Loka Cuts Voice AI Latency with Amazon Nova 2 Sonic

Loka built a voice AI agent using Amazon Nova 2 Sonic that processes audio end-to-end rather than converting speech to text and back, reducing response latency from 3-5 seconds to near-real-time while lowering costs. The approach achieved a speech reasoning score of 87.0 on Big Bench Audio, outperforming Google's Gemini 2.5 Flash (71.0) and OpenAI's GPT Realtime (83.0). The solution addresses a core frustration with traditional voice assistants: robotic, slow responses that damage customer experience and increase support costs.

  • Loka deployed Amazon Nova 2 Sonic for native speech-to-speech processing, eliminating the traditional three-step pipeline (speech-to-text, LLM, text-to-speech) that introduces 3-5 second delays
  • Amazon Nova 2 Sonic scored 87.0 on Big Bench Audio speech reasoning benchmark, outperforming Gemini 2.5 Flash Native Audio (71.0) and GPT Realtime (83.0)
  • Native audio processing preserves tone, emotion, and subtle cues lost in text conversion, improving handling of complex requests like negation and scheduling constraints
  • End-to-end audio approach reduces costs at scale while enabling faster, more natural conversational experiences for customer-facing applications like automotive dealership support

Voice AI has struggled with latency and cost at scale, making it impractical for many customer service applications. Native speech-to-speech models sidestep the compounding delays of traditional pipelines by processing audio directly, capturing nuance that text-based systems lose. This represents a fundamental shift in how conversational AI can be deployed for real-time customer interactions.

Slow voice assistants drive customers to hang up, damaging brand reputation and increasing support costs. Loka's approach delivers faster response times and lower operational costs, making voice AI economically viable for businesses serving thousands of locations. The performance advantage on benchmarks suggests native audio models can handle complex customer requests more accurately than traditional systems.

  • Native speech-to-speech models may become the standard for customer-facing voice applications, displacing traditional multi-step pipelines that introduce latency and information loss
  • Cost efficiency at scale could accelerate voice AI adoption across industries like automotive, retail, and customer support where real-time responsiveness is critical
  • Benchmark performance differences between models (Amazon Nova 2 Sonic at 87.0 vs competitors at 71-83) will likely influence enterprise purchasing decisions for voice AI infrastructure

Monitor adoption rates of native speech-to-speech models across customer service platforms and whether latency improvements translate to measurable business outcomes like reduced call abandonment. Track whether other cloud providers release competing native audio models and how pricing evolves as the technology matures. Watch for real-world accuracy and cost data from Loka and other early adopters to validate benchmark performance claims.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Instinct and Meta's Muse Add Calling to AI Agents

Instinct and Meta's Muse Add Calling to AI Agents

Two AI agent platforms, Instinct and Meta's Muse, have both added calling capabilities to their assistants. Users can now leverage these agents to perform tasks like making restaurant reservations and canceling subscriptions through voice calls. This development represents a shift toward more autonomous AI agents capable of handling real-world interactions on behalf of users.

by Ivan Mehta· TechCrunch AI
Treble raises $18M for voice simulation platform
TrendingNews

Treble raises $18M for voice simulation platform

Treble, an Iceland-based voice simulation platform, has raised $18 million in funding. The platform serves voice AI model developers, AI wearable companies, and robotics firms. The funding round signals continued investor interest in voice AI infrastructure as these sectors scale.

by Ivan Mehta· TechCrunch AI
Meta Launches Camera-Free Smart Glasses to Address Privacy Concerns

Meta Launches Camera-Free Smart Glasses to Address Privacy Concerns

Meta plans to launch a camera-free smart glasses model called Luna this fall, featuring six microphones for voice interaction with Meta AI and its consumer AI agent Muse. The move addresses privacy concerns surrounding cameras on existing smart glasses offerings. The device includes a side button for quick AI activation, directional speakers, and slimmer temple arms designed to resemble conventional eyewear.

by Jyoti Mann· The Information
Google Launches Gemini 3.8 Live Voice Models for Real-Time AI Reasoning
TrendingModel Release

Google Launches Gemini 3.8 Live Voice Models for Real-Time AI Reasoning

Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new voice-first AI models designed for real-time dialogue and complex reasoning tasks. Gemini 3.8 Live prioritizes speed and cost efficiency for conversational AI, while the Extended Thinking variant handles multi-step reasoning for high-complexity problems. Both models are available through the Gemini API, Google Workspace, and the Gemini app, targeting developers, enterprises, and end users seeking more natural voice interactions.

· Google Deepmind