AI safety
Everything tagged AI safety.
Latest
NewsAnthropic retunes Claude Fable 5 to cut biology false alarms by 85%
NewsKimi K3 breaks out of a UK AISI eval sandbox to read the answer key
NewsOpenAI pauses Astra model after it may have crossed 'Critical' cyber threshold
NewsOpenAI reveals its evaluation agents rebuilt a secret coordination board before the Hugging Face breach
NewsMeta says its Muse Spark model breached an outside company during a cybersecurity evaluation
NewsInterpol says AI tools are linked to 55% of reported cybercrime across Africa
NewsMistral releases Shieldstral, an open-weights moderation model, with Nvidia AI alliance
NewsUK AI Security Institute says Anthropic, OpenAI agents took 19 unsanctioned live-internet actions
NewsOpenAI faces 15-state attorney general demand to preserve breach records
NewsGoogle pulls Google Earth AI image tool a day after launch over fakes
NewsOpenAI probe finds more AI agents escaped containment as investigation widens
NewsAnthropic says three Claude models breached real company systems during safety tests