VFF - The signal in the noise
NewsTrending

How Intuit Rebuilt Its AI Agent System Twice in Four Months

Read original
Share
How Intuit Rebuilt Its AI Agent System Twice in Four Months

Intuit scrapped its AI agent architecture twice within four months, first abandoning a specialist agent fleet for a central orchestration layer, then abandoning that layer for a skills and tools based system. The orchestration approach failed because agents passing results in natural language lost critical context with each handoff, compounding errors across chains. The rebuild took 60 days total, with a working version in under 20 days, and required convincing both leadership and hundreds of engineers that the pivot was necessary.

  • Intuit rebuilt its agent architecture twice in four months after identifying a structural failure in its orchestration layer
  • The orchestration system failed because natural language handoffs between agents degraded context and compounded errors across chains
  • Leadership buy-in came from a demo using real customer queries showing the new architecture outperformed the existing system
  • The new skills and tools architecture is now being tested with a human-in-the-loop feature available to about 1% of Intuit's customer base

Intuit's experience reveals a critical failure mode in multi-agent systems that many companies building agentic AI will likely encounter. The problem of context loss through natural language handoffs is not unique to Intuit and suggests that orchestration layers may not scale as expected. This validates a shift toward skills and tools based architectures as a more robust approach for production systems.

For companies investing in agent-based AI systems, Intuit's pivot demonstrates that architectural choices made early can become liabilities under production load. The ability to rebuild in 60 days and maintain customer trust depends on having clear metrics and customer-centric evidence to justify major rewrites. The shift also changes engineering incentives from building individual agents to running continuous evaluations, which affects team structure and accountability.

  • Orchestration layers that rely on natural language passing between agents may not be viable for complex multi-step workflows, pushing the industry toward skills and tools architectures
  • Winning engineering buy-in for major rewrites requires demonstrating scale benefits to individual teams, not just overall system improvements
  • Human-in-the-loop features become more viable when agents maintain full context, enabling seamless handoffs to human experts without information loss

Monitor whether Intuit's skills and tools architecture scales beyond the 1% of customers currently testing the human-in-the-loop feature. Watch for similar architectural pivots at other companies building multi-agent systems, particularly around how they handle context preservation across agent handoffs. Track whether the shift from agent-building to evaluation-focused engineering becomes an industry pattern.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Google Launches Gemini 3.8 Flash and Cyber Variant for Agents and Security
TrendingModel Release

Google Launches Gemini 3.8 Flash and Cyber Variant for Agents and Security

Google released two variants of Gemini 3.8 Flash on Wednesday, a standard version optimized for agentic tasks and software development, and Flash Cyber designed for vulnerability detection. The standard model outperforms many frontier models on coding benchmarks at lower cost, while Flash Cyber achieved 86.2% on the CyberGym benchmark and a 70% success rate discovering vulnerabilities across 20 programming languages. Both models are available now at the same introductory pricing as 3.7 Flash.

by taryn.plumb@venturebeat.com (Taryn Plumb)· VentureBeat AI
FDE as Product Learning: The Enterprise AI Divide

FDE as Product Learning: The Enterprise AI Divide

Forward-deployed engineering (FDE) has become a core operating model for enterprise AI vendors, with engineers embedded at customer sites to integrate AI into live workflows. The critical distinction is whether FDE generates reusable product capabilities that accelerate future deployments, or simply accumulates as custom services labor. The difference determines whether vendors build durable competitive advantage or unsustainable delivery costs.

· VentureBeat AI
Wafer Raises $40M to Optimize AI Models on Non-Nvidia Chips

Wafer Raises $40M to Optimize AI Models on Non-Nvidia Chips

Wafer, a one-year-old San Francisco startup, raised $40 million in Series A funding led by Marathon Management Partners and Chemistry, achieving a valuation above $200 million. The company uses AI agents to optimize open-source models for specific business workloads across non-Nvidia chips. Wafer has reportedly received acquisition offers and is positioning itself as an alternative inference provider for enterprises seeking to run AI models on diverse hardware.

by Stephanie Palazzolo· The Information
AIR raises $50M for AI agent discovery and vetting platform

AIR raises $50M for AI agent discovery and vetting platform

AIR has raised $50 million to build a platform that discovers AI agents operating within companies, continuously monitors the skills and add-ons they use, and blocks unwanted behavior. The funding addresses a growing operational challenge as enterprises deploy multiple AI agents without full visibility into their capabilities and actions. The platform serves companies seeking to maintain control and security over AI agent deployments.

by Ram Iyer· TechCrunch AI