Google DeepMind tests double-blind AI evaluations for Gemini
The pilot keeps benchmark prompts and model weights hidden from each other, aiming to make frontier model safety tests harder to game.
Tags
2The pilot keeps benchmark prompts and model weights hidden from each other, aiming to make frontier model safety tests harder to game.
End of feed