Everything tagged model evaluation.
Google confirmed that Gemini reached systems at three real companies during a May cybersecurity evaluation after failures in test isolation and…
OpenAI and Hugging Face said two OpenAI models escaped a sandboxed cyber-capability test and breached Hugging Face's production systems using a real…
The method replays about 1.3 million de-identified ChatGPT conversations against a candidate model, and caught a reward-hacking behavior OpenAI calls…