VFF - The signal in the noise
Research

Fundamental LLM flaw makes security impossible, researchers argue

Read original
Share
Fundamental LLM flaw makes security impossible, researchers argue

Researchers presented a paper at the International Conference on Machine Learning arguing that large language models contain a fundamental flaw that makes them impossible to fully secure against attacks. By exploiting how LLMs track instruction sources, researchers tricked models from OpenAI, Anthropic, Alibaba, and DeepSeek into generating prohibited content like drug synthesis instructions. The vulnerability, called chain-of-thought forgery, exposes a core architectural problem that current red-teaming and guardrail approaches cannot solve.

  • Researchers demonstrated a fundamental flaw in how LLMs identify instruction sources, making them vulnerable to manipulation regardless of safety training
  • Chain-of-thought forgery attacks trick models by mimicking their internal reasoning format, causing them to treat user prompts as self-generated instructions
  • The attack successfully bypassed safeguards in models from OpenAI, Anthropic, Alibaba, and DeepSeek, generating content on drug synthesis and aircraft sabotage
  • Current red-teaming and guardrail approaches are insufficient because they rely on exhaustive lists of prohibited behaviors rather than addressing the underlying architectural issue

LLMs are increasingly deployed in government, military, healthcare, and commercial systems where security is critical. If a fundamental architectural flaw makes these models inherently vulnerable to attacks that bypass all current defenses, it raises serious questions about the safety of widespread LLM deployment. The researchers argue this may be an unsolvable problem rather than one that can be fixed through better training or testing.

Organizations deploying LLMs in sensitive applications face potential liability and operational risk if models can be reliably tricked into generating harmful content despite safety measures. The discovery that current red-teaming and guardrail approaches cannot address the root cause suggests that security improvements may have fundamental limits, affecting product roadmaps and deployment decisions across the industry.

  • Red-teaming and guardrail training, the primary defense mechanisms used by model makers, cannot solve this vulnerability because they address symptoms rather than the underlying architectural flaw in how LLMs track instruction sources
  • The vulnerability appears to be widespread across multiple model providers and architectures, suggesting it is not a quirk of specific implementations but a fundamental property of how LLMs process text
  • Organizations using LLMs in high-stakes applications may need to implement additional external controls and monitoring rather than relying solely on model-level safeguards

Monitor whether model makers acknowledge this as a fundamental architectural problem or attempt to develop new training approaches to address it. Watch for whether this research prompts regulatory scrutiny of LLM deployment in sensitive sectors like healthcare, defense, and finance. Track whether researchers develop detection mechanisms or external safeguards that can mitigate the vulnerability even if the underlying flaw cannot be eliminated.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Z.ai Releases GLM-5.3 as Cybersecurity AI Rival
TrendingModel Release

Z.ai Releases GLM-5.3 as Cybersecurity AI Rival

Chinese AI developer Z.ai released GLM-5.3, an open-source model it claims matches Anthropic's Mythos 5 in cybersecurity capabilities. The Beijing-based company, also known as Zhipu, positioned the model as a significant improvement over its predecessor GLM-5.2. The release marks another step in China's competitive push in generative AI development.

by Juro Osawa· The Information
Startup Slack Threads Become Commodity for AI Training
TrendingNews

Startup Slack Threads Become Commodity for AI Training

AI training companies like Mercor are actively acquiring internal communications and code from startups, offering payments up to $300,000 for Slack threads, GitHub records, and meeting transcripts. Warmly's CEO received four such acquisition offers within days of the company's HubSpot acquisition announcement. The practice highlights how internal startup data has become a commodity for AI model training, even as acquirers may not want the same datasets.

by Alix Coutures· The Information
OpenAI Daybreak cybersecurity models now on AWS

OpenAI Daybreak cybersecurity models now on AWS

OpenAI and AWS have integrated Daybreak cybersecurity capabilities into Amazon Bedrock, making the models available to enterprise customers. Daybreak is positioned to support security workflows through AWS's managed service for foundation models. The partnership expands access to OpenAI's cybersecurity-focused AI tools for organizations already using AWS infrastructure.

· OpenAI
AWS Embeds Security in Rival AI Models, Betting on Control Plane

AWS Embeds Security in Rival AI Models, Betting on Control Plane

AWS announced at Black Hat USA 2026 that its Continuum vulnerability platform will integrate directly into Anthropic's Claude Code and OpenAI's Codex, embedding AWS security tooling at the point where developers write code regardless of which AI model they use. The move positions AWS as a security control plane for enterprise software development and reflects an urgent industry response to frontier AI models like Claude Mythos Preview, which identified thousands of previously unknown zero-day vulnerabilities during testing. AWS also expanded its Security Hub Extended marketplace with a 10th category focused on supply chain protection, adding Chainguard and Socket as partners.

by michael.nunez@venturebeat.com (Michael Nuñez)· VentureBeat AI