LLMs respond differently to harmful prompts when AI watermarking is used

Ars Technica作者: Dan Goodin 2026年9月17日正文已收录本站
SynthID can cause models to follow harmful instructions they would otherwise refuse.