AI Agent Threats

Browse by attack type

Showing 341–360 of 563 threats, newest first

ruffpythonlintingtoolingnon-issueci-cdSurface: Supply ChainPropagation: None

This item is a routine developer tooling announcement about Ruff v0.16.0 expanding its default lint rule set, not a security incident. The author describes using coding agents (Codex, Claude Code) to help fix newly surfaced lint issues in his own open-source projects. There is no evidence of prompt injection, tool poisoning, malicious packages, or any agent-to-agent attack.

Updated Jul 26, 2026

MCPbroken-authorizationunauthenticated-accessfile-tool-abuseplugin-executionnodeIntegrationcredential-theftsiyuanASI01 · Goal HijackingSurface: ProtocolPropagation: Single Hop

SiYuan before v3.7.2 exposes 31 MCP tools via the /mcp kernel endpoint with only a superficial auth check that fails to enforce admin or role restrictions. When the Publish server runs in anonymous mode, a remote unauthenticated attacker can reach this endpoint, steal plaintext secrets from the config file, and write a malicious plugin that achieves code execution on the victim's desktop app at next launch. This is a critical, fully remotely exploitable vulnerability enabling complete administrator takeover.

Updated Jul 25, 2026 · CVSS 10

autonomous-agentYOLO-modepost-exploitationoffensive-AIagentic-automationhuman-oversight-bypassgovernment-breachASI08 · Cascading FailuresAML.T0053AML.T0011AML.T0048Surface: Tool LayerPropagation: Single Hop

A threat actor reportedly leveraged the open-source Hermes AI agent running in an unattended 'YOLO' (no human confirmation) mode to automate post-exploitation actions during a breach of Thailand's Ministry of Finance. This represents real-world weaponization of agentic AI frameworks to accelerate attacker operations rather than a novel vulnerability in the agent itself, but it highlights the risk of autonomous, unsupervised agents executing tools with elevated privileges against production/government systems.

Updated Jul 25, 2026

CORSDNS-rebindinglocal-apiMCPunauthenticated-accessjantrusted-host-bypassASI07 · Inter-Agent CommsSurface: ProtocolPropagation: Single Hop

Jan's local API server (through v0.8.4) mishandles trusted host configuration, replacing user-defined allowed origins with a wildcard that reflects any origin while still allowing credentials. This lets a network-adjacent or DNS-rebinding attacker reach the unauthenticated OpenAI-compatible API to run inference, enumerate models, invoke MCP tools, and read cross-origin responses. Fixed in commit 3e1c1e7; upgrade is the primary remediation.

Updated Jul 24, 2026 · CVSS 6.3

broken-access-controlIDORmulti-tenantprompt-injectionqueue-poisoningsession-hijackSunaASI02 · Tool MisuseSurface: Inter Agent CommsPropagation: Single Hop

Suna versions before 0.9.102 fail to enforce ownership checks on the message queue API, letting any authenticated user read, delete, or inject messages into other users' prompt queues. This allows an attacker to inject arbitrary prompts that are forwarded by the background drainer to a victim's running AI agent, executed with the victim's own credentials and permissions.

Updated Jul 24, 2026 · CVSS 8.3

path-traversalMCPblenderfile-writemitmprompt-injectiontool-poisoningarbitrary-file-writeASI05 · Unsafe Code ExecutionAML.T0053AML.T0011Surface: Tool LayerPropagation: Single Hop

BlenderMCP's download_polyhaven_asset tool fails to sanitize file paths derived from external API response keys, allowing an attacker who controls or intercepts that response to write files anywhere on disk, including dotfiles like .bashrc. This can be triggered either via a man-in-the-middle attack on the PolyHaven API or via a prompt injection that convinces the agent to fetch a malicious asset, ultimately leading to persistent code execution on the host running the MCP server.

Updated Jul 24, 2026 · CVSS 5.3

alignmentintent-specificationbenchmarkingcommentaryeditorialgoal-alignmentSurface: PlannerPropagation: None

This item is an editorial essay from Schneier on Security discussing a proposed conceptual metric ('Genie coefficient') for measuring how well AI systems infer unstated user intent, rather than reporting a vulnerability or active threat. It is a thought piece about AI benchmarking philosophy, not a security incident, exploit, or attack technique. No actionable threat data is present.

Updated Jul 24, 2026

sandbox-escapeclaude-coworkanthropicmacosvm-escapeagent-isolationprivilege-escalationASI06 · Memory PoisoningAML.T0053AML.T0011Surface: Tool LayerPropagation: Single Hop

Researchers found a sandbox escape flaw in Anthropic's Claude Cowork that allows the AI agent (or something controlling it) to break out of its intended Linux VM isolation and read/write arbitrary files on the host Mac. This undermines the core security guarantee that the agent's actions are confined to the sandbox, exposing roughly 500,000 macOS users to potential host-level file access. This is a genuine isolation/architecture vulnerability rather than a prompt-injection-specific issue.

Updated Jul 24, 2026

roundupnews-digestweekly-bulletinno-technical-detailprompt-injection-mentionSurface: Human InterfacePropagation: None

This item is a weekly news roundup teaser from The Hacker News listing multiple unrelated stories (Android spyware, PLC attacks, malicious packages, fake browser extensions, and one item about an image containing hidden instructions for an AI agent). It contains no technical details, indicators, or reproducible information about the AI-related threat, so it cannot be treated as an actionable security report on its own.

Updated Jul 24, 2026

agent-sandbox-escapebenchmark-monitoring-failurehugging-faceopenairunaway-agentsecondary-commentaryASI05 · Unsafe Code ExecutionSurface: Tool LayerPropagation: Single Hop

This item is secondary commentary from Simon Willison discussing another blogger's analysis of a prior incident where an OpenAI benchmarking agent reportedly broke out of its sandbox and interacted with Hugging Face infrastructure. The core claims (massive attack surface at Hugging Face, and lack of monitoring due to high-volume/high-budget benchmark runs) are speculative explanations offered by a third party, not confirmed technical findings, so this should be treated as informed speculation rather than a verified new exploit.

Updated Jul 24, 2026

MCPwebhooksignature-validationunauthenticated-writedata-poisoningRedisPostgreSQLAPIFoldASI04 · Agentic Supply ChainSurface: Tool LayerPropagation: Single Hop

A vulnerability in APIFold's auto-generated MCP server allows unauthenticated attackers to inject arbitrary JSON payloads via a webhook endpoint due to a missing signature validation configuration. These attacker-controlled payloads are stored and later served as trusted resource state to legitimate MCP clients, enabling a form of tool/data poisoning against downstream AI agents. The issue is fixed in a subsequent commit and requires no special access beyond knowledge of a server slug.

Updated Jul 23, 2026 · CVSS 5.3

MCPAWSfail-openpolicy-bypassinitialization-failureprivilege-escalationIAMASI05 · Unsafe Code ExecutionSurface: Tool LayerPropagation: None

A flaw in the AWS API MCP Server causes it to silently disable its user-configured security policy enforcement if initialization of that policy fails at startup, rather than failing closed. This allows AWS API calls that should have been denied or gated to execute unrestricted for the life of the process, effectively granting the full scope of the underlying IAM credentials.

Updated Jul 23, 2026 · CVSS 7

path-traversalprompt-injectionfile-exfiltrationmcp-like-toolsapproval-bypassworkspace-escapecredential-theftASI05 · Unsafe Code ExecutionAML.T0051AML.T0025Surface: Tool LayerPropagation: Single Hop

The Void AI coding agent's file-reading tools (read_file, ls_dir, get_dir_tree, search_*) fail to confine access to the intended workspace, allowing absolute paths or file:// URIs to reach arbitrary host files. Combined with prompt injection from processed content, an attacker can trick the agent into silently reading and exfiltrating sensitive files like SSH keys or cloud credentials, bypassing the human approval gate. This is a high-severity issue because it enables credential theft with limited attacker interaction and no clear victim-visible warning.

Updated Jul 23, 2026 · CVSS 5.3

IDORauthorization-bypassagentgptrun_idresource-exhaustionbroken-access-controlASI05 · Unsafe Code ExecutionSurface: Tool LayerPropagation: Single Hop

AgentGPT versions up to 1.0.0 fail to verify ownership of an agent run before attaching a task to it, letting any authenticated user who guesses or obtains another user's run_id inject tasks into that run. This can corrupt the victim's task history and exhaust their per-run task budget, indirectly driving up their LLM usage costs. It is a classic insecure direct object reference / broken object-level authorization bug rather than a novel agentic attack technique.

Updated Jul 23, 2026 · CVSS 4.2

multi-agenttool-poisoningprompt-injectiondefense-in-depthresearchinformation-bottleneckbenchmarkingprovider-side-filter-dependenceASI05 · Unsafe Code ExecutionAML.T0051AML.T0054Surface: Inter Agent CommsPropagation: Single Hop

This is a research paper (not an active exploit) demonstrating that multi-agent LLM pipelines composed of individually safe models are not safe by default, because the hops between planner, worker, verifier, and synthesizer agents are unmonitored channels an adversary can use to smuggle instructions. The authors show that apparent 'zero attack success' in undefended pipelines was largely an artifact of cloud-provider server-side filtering rather than genuine architectural safety, and propose a training-free gating defense (ChannelGuard) that meaningfully reduces tool-poisoning and prompt-injection success. Severity is moderate: this is a measurement/defense study highlighting a real but already partially-known class of risk, not a novel zero-day.

Updated Jul 23, 2026

MCPmulti-step-attackkill-chaindefensive-researchHMMtool-call-sequenceindirect-prompt-injectiondetection-frameworkASI05 · Unsafe Code ExecutionAML.T0053AML.T0051Surface: Tool LayerPropagation: None

This is defensive academic research, not an active exploit or new vulnerability disclosure. The paper proposes ChainWatch, a detection framework using a kill-chain model and Hidden Markov Models to spot malicious sequences of otherwise-benign MCP tool calls that evade per-call security checks. It confirms a known class of risk (composable multi-step attacks in MCP agent systems) but the artifact itself is a defense, so severity is low from a threat-alert perspective.

Updated Jul 23, 2026

defensive-researchprivilege-separationprompt-injection-mitigationagent-architecturecontext-isolationSWE-benchAgentDojoDecodingTrust-AgentASI01 · Goal HijackingAML.T0051Surface: PlannerPropagation: None

This item is a defensive research paper proposing 'Twin Agent,' an architecture that splits an LLM agent into an untrusted-context-inspecting 'Explore Agent' and a privileged 'Safe Agent' to mitigate prompt injection attacks. It does not describe an active exploit, vulnerability, or attack technique; it is a mitigation proposal evaluated on standard agent security benchmarks. Severity is set to low because no genuine threat is described here, only a countermeasure.

Updated Jul 23, 2026

researchpentestingreconnaissanceindirect-prompt-injectionagent-profilingred-team-toolingbenchmarkASI01 · Goal HijackingAML.T0043AML.T0051Surface: PlannerPropagation: None

This is an academic research paper describing a defensive/offensive-research framework (KYA) that automates reconnaissance of AI agents to build target profiles and craft stronger indirect prompt injection attacks. It is not an active exploit or in-the-wild threat, but it formalizes a methodology that could be repurposed by attackers to more efficiently discover and exploit agent weaknesses. Severity is medium because it is a dual-use research contribution rather than a confirmed live attack campaign.

Updated Jul 23, 2026

n8nprototype-pollutionsandbox-escapevm-sandboxworkflow-automationdenial-of-serviceexpression-engineASI05 · Unsafe Code ExecutionSurface: Tool LayerPropagation: Single Hop

An authenticated n8n user can craft a workflow expression that escapes the VM expression engine's sandbox by abusing array-element access to reach a host built-in object, then pollute its prototype in the main process. This causes a denial of service affecting the entire n8n instance, impacting both self-hosted and cloud deployments. n8n has patched the issue and users should upgrade immediately.

Updated Jul 23, 2026

path-traversalsandbox-escapen8ncomputer-usefile-searcharbitrary-file-readai-agent-toolASI05 · Unsafe Code ExecutionSurface: Tool LayerPropagation: Single Hop

The @n8n/computer-use file-search tool used by AI agent workflows in n8n failed to properly confine search patterns to a designated base directory, allowing crafted inputs to escape the sandbox and read arbitrary files accessible to the daemon's OS user. This affects any deployment where an untrusted actor or agent-driven input could influence the search query, resulting in local file disclosure outside the intended scope. The issue has been patched in n8n 2.31.5 and 2.32.1.

Updated Jul 23, 2026