Have the frontier labs mixed up AI safety and security?

Martin Alderson: “Note the phrasing - ‘we have largely solved the threat of prompt injection in practice’. Now look at the benchmark attached to that very tweet - it’s nowhere near solved. The Opus 5 score (the best score) - fails to a prompt injection attack 2% of the time with 15 attempts.”

[bookmark]