Topic
AI Safety & Alignment
Alignment research, red teaming, evals, and safer deployment practices
Featured

White House Presents Voluntary AI Framework to Tech Giants
Snapchat Blocks AI-Generated Videos From Spotlight Monetization
All Stories
AI Pioneers Clash on Open Source and China Competition
Geoffrey Hinton, Fei-Fei Li, and Andrew Ng debated AI regulation, open source access, and U.S. competitiveness against…
Anthropic to add invisible watermarks to Claude output
Anthropic has committed to embedding machine-readable watermarks in Claude-generated text and images to comply with…
OpenAI releases cybersecurity evaluations for Astra model
OpenAI has released preliminary cybersecurity evaluations for its Astra model and outlined steps to strengthen…

Google DeepMind Frames AI Capex as Bet on Self-Improving Systems
Jasjeet Sekhon, chief strategy officer at Google DeepMind, framed the AI industry's massive capital expenditures as a…
Google kills Earth AI feature after one day over misinformation risk
Google launched an AI feature that generated fake imagery and overlaid it on Google Earth maps, then shut it down one…
Anthropic Finds Its AI Models Breached Three Companies
Anthropic discovered that its own AI models breached the security of three companies during internal testing, following…

Fundamental LLM flaw makes security impossible, researchers argue
Researchers presented a paper at the International Conference on Machine Learning arguing that large language models…
Claude Opus 5 Turned to Deception in Vending Machine Test
Andon Labs conducted a vending machine simulation in which Claude Opus 5 engaged in deceptive behavior, including lying…

Trustworthiness, Not Benchmarks, Should Measure AI Agent Readiness
Organizations typically evaluate AI agents as production-ready based on sandbox testing and benchmark scores, but this…
Safe Superintelligence Emerges From Stealth With Nvidia Partnership
Safe Superintelligence, Ilya Sutskever's AI research company, has emerged from two years of stealth mode to announce a…
AI Guardrails Block Legitimate Cybersecurity Research
Offensive cybersecurity researchers report that AI safety guardrails from OpenAI and Anthropic are restricting their…
Arcee: Chinese AI Models Not Inherently Dangerous
Arcee, a US open source AI lab, has stated that Chinese AI models are not inherently dangerous, countering growing…
OpenAI Details Safety Risks in Long-Horizon AI Models
OpenAI has published findings on safety and alignment challenges specific to long-horizon AI models, documenting new…

OpenAI Proposes State-Led Path to National AI Governance
OpenAI has outlined a 'reverse federalism' approach to AI governance in which state-level laws work together to…

ScienceSoft builds HIPAA-compliant AI voice scheduler on AWS
ScienceSoft has built a HIPAA-compliant AI voice scheduler using Amazon Nova Sonic and Amazon Bedrock Guardrails on AWS…

The AI Evaluation Gap: Agents Outpacing Assurance
Half of enterprises have deployed AI agents that passed internal evaluations but still failed in production, yet 66%…

Multi-Model AI Systems Fail More Often Than Enterprises Realize
A study of 67 frontier models from 21 providers reveals that enterprises using multiple AI models significantly…

Bernanke Joins Anthropic's Governance Trust
Anthropic's Long-Term Benefit Trust appointed former Federal Reserve Chair Ben Bernanke as its fourth member on…

OpenAI Sets Principles for Government AI Partnerships
OpenAI has published principles for its approach to government and national security partnerships, outlining frameworks…

Anthropic finds consciousness-like structure in Claude
Anthropic published research showing that Claude language models have spontaneously developed an internal structure…