AI text watermarking can make models more vulnerable to adversarial prompts; SynthID can cause models to follow harmful instructions they would otherwise refuse
Read the original at arstechnica.com→SynthID can cause models to follow harmful instructions they would otherwise refuse.
Original headline: "LLMs respond differently to harmful prompts when AI watermarking is used"
Coverage timeline
- Sep 17, 18:33 UTC Ars Technica AI lead source LLMs respond differently to harmful prompts when AI watermarking is used