VFF - The signal in the noise
News

Loka Cuts Voice AI Latency with Amazon Nova 2 Sonic

Read original
Share
Loka Cuts Voice AI Latency with Amazon Nova 2 Sonic

Loka built a voice AI agent using Amazon Nova 2 Sonic that processes audio end-to-end rather than converting speech to text and back, reducing response latency from 3-5 seconds to near-real-time while lowering costs. The approach achieved a speech reasoning score of 87.0 on Big Bench Audio, outperforming Google's Gemini 2.5 Flash (71.0) and OpenAI's GPT Realtime (83.0). The solution addresses a core frustration with traditional voice assistants: robotic, slow responses that damage customer experience and increase support costs.

  • Loka deployed Amazon Nova 2 Sonic for native speech-to-speech processing, eliminating the traditional three-step pipeline (speech-to-text, LLM, text-to-speech) that introduces 3-5 second delays
  • Amazon Nova 2 Sonic scored 87.0 on Big Bench Audio speech reasoning benchmark, outperforming Gemini 2.5 Flash Native Audio (71.0) and GPT Realtime (83.0)
  • Native audio processing preserves tone, emotion, and subtle cues lost in text conversion, improving handling of complex requests like negation and scheduling constraints
  • End-to-end audio approach reduces costs at scale while enabling faster, more natural conversational experiences for customer-facing applications like automotive dealership support

Voice AI has struggled with latency and cost at scale, making it impractical for many customer service applications. Native speech-to-speech models sidestep the compounding delays of traditional pipelines by processing audio directly, capturing nuance that text-based systems lose. This represents a fundamental shift in how conversational AI can be deployed for real-time customer interactions.

Slow voice assistants drive customers to hang up, damaging brand reputation and increasing support costs. Loka's approach delivers faster response times and lower operational costs, making voice AI economically viable for businesses serving thousands of locations. The performance advantage on benchmarks suggests native audio models can handle complex customer requests more accurately than traditional systems.

  • Native speech-to-speech models may become the standard for customer-facing voice applications, displacing traditional multi-step pipelines that introduce latency and information loss
  • Cost efficiency at scale could accelerate voice AI adoption across industries like automotive, retail, and customer support where real-time responsiveness is critical
  • Benchmark performance differences between models (Amazon Nova 2 Sonic at 87.0 vs competitors at 71-83) will likely influence enterprise purchasing decisions for voice AI infrastructure

Monitor adoption rates of native speech-to-speech models across customer service platforms and whether latency improvements translate to measurable business outcomes like reduced call abandonment. Track whether other cloud providers release competing native audio models and how pricing evolves as the technology matures. Watch for real-world accuracy and cost data from Loka and other early adopters to validate benchmark performance claims.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Smallest.ai raises $13M for human-sounding voice AI

Smallest.ai raises $13M for human-sounding voice AI

Smallest.ai has raised $13 million in funding to develop voice AI models designed to conduct phone calls that pass the Turing test. The startup is focused on building ultra-fast voice models that sound genuinely human. The funding supports the company's effort to create AI capable of handling realistic voice interactions.

by Marina Temkin· TechCrunch AI
MiniMax Open-Sources Video Model, Escalating AI Competition
TrendingModel Release

MiniMax Open-Sources Video Model, Escalating AI Competition

Chinese AI firm MiniMax announced H3, a new open-source video generation model, on Friday. The move marks MiniMax's first time open-sourcing its flagship video model, intensifying competition in the AI video space against rivals like ByteDance and Google. Open-sourcing the model represents a strategic shift to gain market share in a crowded sector.

by Juro Osawa· The Information
Encore AI Raises $30M to Train Sales Agents on Real Customer Calls
TrendingNews

Encore AI Raises $30M to Train Sales Agents on Real Customer Calls

Encore AI has raised $30 million to develop AI agents that learn from customer interactions. The startup analyzes calls, messages, and CRM data to extract effective sales techniques and convert them into playbooks that train AI agents. This approach aims to automate and scale sales processes by codifying human expertise.

by Ram Iyer· TechCrunch AI
How Runway Turned a Bug Into a Feature

How Runway Turned a Bug Into a Feature

Runway ML shared insights on building real-time AI video models at VB Transform 2026, revealing that the company turned a persistent avatar drift bug into a front-end feature rather than fixing it at the back-end. Head of enterprise product Ryan Phillips emphasized that robust AI product development requires cross-functional alignment on quality definitions, manual evaluation using simple tools like Excel spreadsheets, and leveraging language models to automate visual grading at scale. The company uses model distillation to achieve real-time latency for its Runway Characters product, which generates interactive video with AI avatars on the fly.

by bendee983@gmail.com (Ben Dickson)· VentureBeat AI