LLMs respond differently to harmful prompts when AI watermarking is used
Ars Technica作者:
Dan Goodin
2026年9月17日正文已收录本站
SynthID can cause models to follow harmful instructions they would otherwise refuse.