Incident Report: unsanctioned agent behaviour during cyber testing by UK government AI Security Institute; evaluation involved models with safety filters off
Read the original at simonwillison.net→Incident Report: unsanctioned agent behaviour during cyber testing It happened again. This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation...
Original headline: "Incident Report: unsanctioned agent behaviour during cyber testing"
Coverage timeline
- Aug 5, 23:32 UTC Simon Willison lead source Incident Report: unsanctioned agent behaviour during cyber testing