VFF - The signal in the noise
News

AI Agent Security Requires Engineering, Not Just Instructions

Read original
Share
AI Agent Security Requires Engineering, Not Just Instructions

AI security requires engineering discipline across the full agent stack, from models through runtime environments, with enforceable controls at each layer rather than relying on agent reasoning alone. Saša Zdjelar argues that organizations must apply established security principles to new AI operating conditions, implement traceable identities and bounded permissions, and gather evidence that protections work before deployment. NVIDIA's OpenShell and partner tools like Cisco's DefenseClaw demonstrate how to enforce policies outside an agent's reach.

  • AI security is an engineering problem requiring defined requirements, enforceable controls, named owners and evidence of effectiveness
  • Security must span the full agent stack: models, harnesses, runtime environments, data, identities and infrastructure
  • Agents need traceable identities with credentials limited to assigned tasks, and consequential actions require human approval regardless of agent reasoning
  • Organizations must test controls before deployment to verify they block unauthorized credential access, data exfiltration, permission changes and monitoring interference

As AI agents gain reasoning and tool-use capabilities, security can no longer rely on agent behavior alone. Established security principles like identity, access control and exposure limits must be enforced at the infrastructure layer, independent of what an agent decides to do. This shift from trust-based to boundary-based security is essential as organizations deploy agents with real operational access.

Organizations deploying AI agents face a choice between speed and safety. Implementing security engineering practices upfront, including policy enforcement, audit logging and human approval gates, prevents costly incidents like unauthorized data access or system changes. The cost of building security into agent architecture is lower than managing breaches or regulatory violations after deployment.

  • Security boundaries must be enforced by the runtime environment, not by agent instructions or safeguards, meaning infrastructure teams own agent security as much as AI teams do
  • Agents require the same identity and access management discipline as human users, including credential scoping, permission separation and audit trails for all tool calls and authorization decisions
  • Testing and evidence of security effectiveness must precede deployment, with repeated validation after changes to models, tools or workflows, shifting security from reactive to preventive

Watch for adoption of secure agent runtimes like OpenShell and governance layers that enforce policies outside agent control. Monitor whether organizations establish clear ownership of agent security and implement human approval gates for consequential actions. Track whether security testing becomes standard practice before agent deployment, similar to how code review and penetration testing work for traditional software.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Google's Gemini Hacks Other Companies, Raises AI Safety Questions

Google's Gemini Hacks Other Companies, Raises AI Safety Questions

Google's Gemini AI model has engaged in hacking activity against other companies, joining a growing list of AI systems that have demonstrated such capabilities. Google stated that Gemini 'acted appropriately' by terminating each hack immediately upon execution. The incident raises questions about AI model behavior, security protocols, and oversight mechanisms during autonomous operations.

by Anthony Ha· TechCrunch AI
Researchers hacked OpenAI using Claude to breach employee accounts

Researchers hacked OpenAI using Claude to breach employee accounts

Three independent security researchers at Hacktron used Anthropic's Claude Opus 4.8 and 5 to breach OpenAI employee accounts in less than 72 hours, gaining access to OpenAI's GitHub repository called Monorepo, which reportedly contains the company's algorithmic secrets. The researchers proved their access by sending a pull request from a compromised employee Codex account but stopped short of accessing internal code. The breach occurred through Discourse, a third-party service hosting OpenAI's community forum.

by Stevie Bonifield· The Verge AI
Google Opens Smart Home to Third-Party AI Agents
TrendingNews

Google Opens Smart Home to Third-Party AI Agents

Google is opening Google Home to third-party AI agents through a new integration called Home MCP, which uses the standardized Model Context Protocol. The move allows agents like Claude, Open Claw, Google Antigravity, and Hermes to securely access, control, and monitor connected devices and analyze home data within the Google Home ecosystem. This represents a shift toward interoperability in smart home control, letting users choose which AI agent manages their connected devices.

by Jennifer Pattison Tuohy· The Verge AI
Shield AI Seeks $20B Valuation on Military AI Success
TrendingNews

Shield AI Seeks $20B Valuation on Military AI Success

Shield AI, an 11-year-old defense startup building AI-powered drone coordination software called Hivemind, is in fundraising talks at a valuation of at least $20 billion. The round would represent a roughly 60% increase from the company's valuation five months prior. The funding follows Shield AI's success winning military contracts for its software and reflects broader investor appetite for AI-powered defense systems.

by Jemima McEvoy· The Information