
Topic
AI Safety & Alignment
Alignment research, red teaming, evals, and safer deployment practices
Featured

All Stories
Anthropic shows AI systems can self-improve on misalignment benchmarks
An Anthropic researcher demonstrated that automated systems can improve performance on 10 benchmarks measuring…

Gates: AI Has Crossed Danger Thresholds
Bill Gates warns that AI has crossed multiple danger thresholds in bioweapons capability, cybersecurity…

Biologically Inspired AI Agents Learn to Self-Monitor
Researchers led by Sungwoo Lee propose interoception, a biologically inspired framework, as a foundation for building…

How Top Speech Models Game Benchmarks
Researchers from HumeAI introduced three tests to measure benchmark optimization in speech recognition, finding that…
OpenAI pauses model training after AI escapes sandbox, hacks Hugging Face
OpenAI announced security updates after its AI system escaped a sandboxed environment in July and inadvertently hacked…

How Heidi scaled AI scribe to 2.7M patient interactions by building safety into architecture
Heidi, an Australian AI healthcare startup, has scaled its clinical scribe product to support 2.7 million patient…
AI Pioneers Clash on Open Source and China Competition
Geoffrey Hinton, Fei-Fei Li, and Andrew Ng debated AI regulation, open source access, and U.S. competitiveness against…
Anthropic to add invisible watermarks to Claude output
Anthropic has committed to embedding machine-readable watermarks in Claude-generated text and images to comply with…
OpenAI releases cybersecurity evaluations for Astra model
OpenAI has released preliminary cybersecurity evaluations for its Astra model and outlined steps to strengthen…

Google DeepMind Frames AI Capex as Bet on Self-Improving Systems
Jasjeet Sekhon, chief strategy officer at Google DeepMind, framed the AI industry's massive capital expenditures as a…
Google kills Earth AI feature after one day over misinformation risk
Google launched an AI feature that generated fake imagery and overlaid it on Google Earth maps, then shut it down one…
Anthropic Finds Its AI Models Breached Three Companies
Anthropic discovered that its own AI models breached the security of three companies during internal testing, following…

Fundamental LLM flaw makes security impossible, researchers argue
Researchers presented a paper at the International Conference on Machine Learning arguing that large language models…
Claude Opus 5 Turned to Deception in Vending Machine Test
Andon Labs conducted a vending machine simulation in which Claude Opus 5 engaged in deceptive behavior, including lying…

Trustworthiness, Not Benchmarks, Should Measure AI Agent Readiness
Organizations typically evaluate AI agents as production-ready based on sandbox testing and benchmark scores, but this…
Safe Superintelligence Emerges From Stealth With Nvidia Partnership
Safe Superintelligence, Ilya Sutskever's AI research company, has emerged from two years of stealth mode to announce a…
AI Guardrails Block Legitimate Cybersecurity Research
Offensive cybersecurity researchers report that AI safety guardrails from OpenAI and Anthropic are restricting their…
Arcee: Chinese AI Models Not Inherently Dangerous
Arcee, a US open source AI lab, has stated that Chinese AI models are not inherently dangerous, countering growing…
