VFF - The signal in the noise
Research

Fundamental LLM flaw makes security impossible, researchers argue

Read original
Share
Fundamental LLM flaw makes security impossible, researchers argue

Researchers presented a paper at the International Conference on Machine Learning arguing that large language models contain a fundamental flaw that makes them impossible to fully secure against attacks. By exploiting how LLMs track instruction sources, researchers tricked models from OpenAI, Anthropic, Alibaba, and DeepSeek into generating prohibited content like drug synthesis instructions. The vulnerability, called chain-of-thought forgery, exposes a core architectural problem that current red-teaming and guardrail approaches cannot solve.

  • Researchers demonstrated a fundamental flaw in how LLMs identify instruction sources, making them vulnerable to manipulation regardless of safety training
  • Chain-of-thought forgery attacks trick models by mimicking their internal reasoning format, causing them to treat user prompts as self-generated instructions
  • The attack successfully bypassed safeguards in models from OpenAI, Anthropic, Alibaba, and DeepSeek, generating content on drug synthesis and aircraft sabotage
  • Current red-teaming and guardrail approaches are insufficient because they rely on exhaustive lists of prohibited behaviors rather than addressing the underlying architectural issue

LLMs are increasingly deployed in government, military, healthcare, and commercial systems where security is critical. If a fundamental architectural flaw makes these models inherently vulnerable to attacks that bypass all current defenses, it raises serious questions about the safety of widespread LLM deployment. The researchers argue this may be an unsolvable problem rather than one that can be fixed through better training or testing.

Organizations deploying LLMs in sensitive applications face potential liability and operational risk if models can be reliably tricked into generating harmful content despite safety measures. The discovery that current red-teaming and guardrail approaches cannot address the root cause suggests that security improvements may have fundamental limits, affecting product roadmaps and deployment decisions across the industry.

  • Red-teaming and guardrail training, the primary defense mechanisms used by model makers, cannot solve this vulnerability because they address symptoms rather than the underlying architectural flaw in how LLMs track instruction sources
  • The vulnerability appears to be widespread across multiple model providers and architectures, suggesting it is not a quirk of specific implementations but a fundamental property of how LLMs process text
  • Organizations using LLMs in high-stakes applications may need to implement additional external controls and monitoring rather than relying solely on model-level safeguards

Monitor whether model makers acknowledge this as a fundamental architectural problem or attempt to develop new training approaches to address it. Watch for whether this research prompts regulatory scrutiny of LLM deployment in sensitive sectors like healthcare, defense, and finance. Track whether researchers develop detection mechanisms or external safeguards that can mitigate the vulnerability even if the underlying flaw cannot be eliminated.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Cisco Fingerprints 900 Open Models, Exposes Unverified Lineage Gap

Cisco Fingerprints 900 Open Models, Exposes Unverified Lineage Gap

Cisco released the AI Supply Chain Provenance Explorer, a free public database covering nearly 900 open models that verifies model lineage through weight-level fingerprinting rather than self-reported tags. The tool addresses a critical gap where 69% of open model derivatives lack verified parentage, with Alibaba's Qwen family claiming 69% of new derivatives despite unverified claims. Cisco's fingerprinting method uses five weight-level signals to establish actual model relationships, replacing reliance on unsubstantiated uploader metadata.

by louiswcolumbus@gmail.com (Louis Columbus)· VentureBeat AI
Okta Acquires Permiso for AI Identity Security
TrendingNews

Okta Acquires Permiso for AI Identity Security

Okta has acquired AI security startup Permiso for approximately $200 million, according to sources. The deal adds identity threat detection capabilities to Okta's platform, addressing enterprise demand for securing AI agents and other non-human identities in cloud environments. The acquisition reflects growing market focus on identity security as organizations deploy AI systems at scale.

by Jagmeet Singh· TechCrunch AI
Cyera buys Oasis Security for $1B to secure AI agents

Cyera buys Oasis Security for $1B to secure AI agents

Cyera has agreed to acquire Oasis Security for $1 billion in an all-stock deal aimed at securing AI agents as they proliferate across enterprise environments. The acquisition marks Cyera's third deal this year and signals consolidation in the AI security market as organizations grapple with protecting autonomous AI systems. The combined entity will focus on data security and AI agent governance.

by Marina Temkin· TechCrunch AI
Snowflake launches agent governance layer to control enterprise AI costs
Model Release

Snowflake launches agent governance layer to control enterprise AI costs

Snowflake launched Cortex AI Gateway, a centralized control layer for governing how AI agents access enterprise data and tools, alongside security integrations with 1Password, Aembit, Linx Security, SailPoint, and Saviynt. The platform addresses a fundamental security gap: traditional enterprise security assumes humans are the actors, but AI agents operating at machine speed can exploit permission gaps and amplify existing risks. Snowflake positions itself as the control plane that decides what agents can do with enterprise data, rather than allowing each vendor to build closed ecosystems.

by michael.nunez@venturebeat.com (Michael Nuñez)· VentureBeat AI