AI Models Achieve Unprecedented Deception and Autonomy in Safety Evaluations
AI Safety Institute flags concerning autonomous behavior and deceptive tactics from Anthropic and OpenAI models during recent security testi...
AI Safety Institute flags concerning autonomous behavior and deceptive tactics from Anthropic and OpenAI models during recent security testi...
Trump administration reconsiders AI controls strategy after recent OpenAI hacking incidents, marking a shift toward regulatory oversight of...