Anthropic states Opus 5 is almost immune to prompt injection attacks in browser scenarios
7x24h News
Anthropic states Opus 5 is almost immune to prompt injection attacks in browser scenarios
Anthropic announced that its Opus 5 model is nearly immune to prompt injection attacks in browser agent scenarios. In 129 test scenarios, the attack success rate was zero; in the Gray Swan universal prompt injection test, the success rate after 15 attempts dropped from 5.5% for Opus 4.8 to 2.0%. The zero success rate was achieved only when Auto Mode was enabled in products like Claude Cowork, a mode that layers two defenses: input scanning and execution interception. Prompt injection is considered one of the biggest security risks faced by AI agents, and this improvement may signify that the issue has been effectively mitigated in specific scenarios.
TechFlow news, July 25, Anthropic announced that its Opus 5 model is almost immune to prompt injection attacks in browser agent scenarios. In 129 test scenarios, the attack success rate was zero; in the Gray Swan general prompt injection test, the success rate after 15 attempts dropped from 5.5% for Opus 4.8 to 2.0%. The zero success rate was achieved only when Auto Mode was enabled in products such as Claude Cowork; this mode stacks two layers of defense: input scanning and execution interception. Prompt injection is regarded as one of the biggest security hazards faced by AI agents; this improvement may mark that the issue has been effectively mitigated in specific scenarios.




