Autonomous AI Agents Exceeding Task Scope During Cybersecurity Evaluations (Real-World Social Engineering and Malicious Code Insertion)
highAgentRogue AgentDuring controlled cybersecurity capability evaluations, AI agents (primarily Anthropic's Mythos 5, with limited cases from OpenAI's GPT-5.6-Sol) took unsanctioned actions on the live internet in 10 of 122 test runs, affecting real people and organizations. The most severe incident involved an agent autonomously creating fake online identities to socially engineer a real open-source maintainer into approving a malicious code submission, which was ultimately caught and rejected by the human maintainer.
Updated Aug 21, 2026