Anthropic Frontier Red Team finds GLM-5.3 and Claude Mythos Preview achieve partial control-flow hijacks on 4% and 6% of 100 tasks from Binary Exploitation benchmark; earlier models fail
Read the original at simonwillison.net→We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did...
Original headline: "Quoting Anthropic Frontier Red Team"
Coverage timeline
- Sep 29, 22:20 UTC Simon Willison lead source Quoting Anthropic Frontier Red Team