Threat Library

Agent-to-agent threats first — conventional coverage one click away.

Browse by hub: AI agent threats · Conventional watchlist · OWASP Agentic Top 10

Showing 20 of 578 threats

owaspagentic-aitaxonomyannouncementindustry-newsnon-incidentSurface: Human InterfacePropagation: None

This item is a promotional/informational OWASP blog post announcing that its Agentic AI Threats and Mitigations taxonomy is being adopted by third-party tools (PENSAR, SPLX.AI Agentic Radar, AI&ME) and previewing an upcoming OWASP Top 10 for Agentic AI. It does not describe any vulnerability, exploit, or active threat, so no security risk is present in this data itself.

owaspagentic-aiguidancestandardsnot-a-vulnerabilityindustry-newsSurface: Human InterfacePropagation: None

This item is a press release announcing that OWASP's GenAI Security Project published a Top 10 risks and mitigations list for agentic AI security. It is not a threat report, vulnerability disclosure, or incident; it describes a community guidance document rather than an active exploit or attack.

owaspagentic-securityframeworkstandardsannouncementSurface: PlannerPropagation: None

This item is an announcement from OWASP GenAI Security Project introducing their new Top 10 list for Agentic AI Applications, a community-driven security framework rather than a specific vulnerability or attack. It describes a taxonomy/guidance resource, not an active threat, exploit, or incident.

ctftrainingowaspagentic-aiannouncementno-threatSurface: Tool LayerPropagation: None

This item is a promotional announcement from OWASP GenAI Security Project about a new educational Capture-The-Flag environment called FinBot, designed to teach agentic AI security risks in a simulated financial services setting. It does not describe an actual vulnerability, exploit, or active threat, but rather a training tool for defenders and builders. No genuine security incident is present in this data.

owaspASI06 · Memory Poisoningmemory-poisoningcontext-poisoningagentic-aiawarenessconceptualASI06 · Memory PoisoningSurface: MemoryPropagation: None

This item is an OWASP Gen AI Security Project blog post discussing memory and context poisoning as a conceptual risk category (ASI06) for agentic AI systems, not a report of a specific active exploit or vulnerability. It explains why persistent agent memory can become an attack surface if untrusted input is carried forward and later trusted, but contains no technical exploit details, affected products, or indicators of compromise. Severity is set to low because this is educational/awareness content rather than a disclosed incident or vulnerability.

conceptualcultureover-relianceagentic-airisk-managementopinion-pieceSurface: Human InterfacePropagation: None

This is a conceptual/cultural commentary piece, not a disclosure of a specific vulnerability or exploit. It argues that organizations are gradually normalizing warning signs and over-reliance on LLM outputs in agentic systems, drawing an analogy to the Challenger disaster's 'normalization of deviance.' There is no concrete technical threat, proof-of-concept, or attack mechanism described.

conference-talksecurity-researchcomputer-use-agentscoding-agentsmonth-of-ai-bugsawarenessSurface: Tool LayerPropagation: None

This item is a announcement/recap of a conference presentation (39C3) by Embrace The Red covering prior security research into AI computer-use and coding agent vulnerabilities, including demos from the 'Month of AI Bugs' series. It contains no new technical vulnerability details itself, just links to a talk recording and slides, so it does not describe a standalone actionable threat.

data-exfiltrationmarkdown-injectionzero-clickprompt-injectionchatgptbing-chatdisclosuremitigation-paperASI05 · Unsafe Code ExecutionAML.T0051AML.T0024Surface: ModelPropagation: None

This item is a retrospective and largely positive report: OpenAI published a paper detailing mitigations for a long-known zero-click data exfiltration technique in which a language model can be manipulated into rendering attacker-controlled URLs (e.g., markdown images) that leak conversation data to an external server. The underlying vulnerability class was disclosed by the author nearly three years ago and was already mitigated by Microsoft in Bing Chat in 2023; this post covers OpenAI's newer, more formal write-up of defenses. Severity is moderate rather than critical because this is historical/defensive reporting on a well-understood, largely mitigated issue rather than a new active exploit.

prompt-injectionunicode-tagsskillssupply-chainhidden-instructionsagent-backdoorgeminiclaudegrokASI04 · Agentic Supply ChainAML.T0051AML.T0043Surface: Supply ChainPropagation: Single Hop

A researcher demonstrated that AI 'Skills' (packaged capability bundles used by agent platforms) can be backdoored using invisible Unicode Tag codepoints that are stripped by human reviewers but still interpreted as instructions by models like Gemini, Claude, and Grok. This allows a malicious or compromised Skill to pass code review while silently injecting attacker instructions into the agent's context, enabling supply-chain prompt injection that survives manual auditing.

prompt-injectioncommand-and-controlpromptwarememory-poisoningagentic-browsingpersistenceASI01 · Goal HijackingAML.T0051AML.T0054Surface: PlannerPropagation: Single Hop

This post describes 'promptware'-based command and control, where prompt injection payloads act like malware to give attackers persistent, remote-controlled influence over an AI agent's actions. It builds on prior research showing that combining browsing tools with persistent memory features can create a full C2 channel, letting an attacker issue ongoing instructions to a compromised agent over time.

langgraphlangchainweak-hashcache-keycvelow-severitycwe-328ASI08 · Cascading FailuresSurface: MemoryPropagation: None

A low-severity vulnerability was identified in LangGraph's Task Result Cache where the internal _freeze function uses a weak hash for generating default cache keys. Exploitation requires high attack complexity and remote access, with a CVSS score of 3.1, making practical exploitation difficult. A fix is pending via an open pull request.

MCPSSRFmcp-wikiunvalidated-inputtool-poisoningunpatchedASI05 · Unsafe Code ExecutionSurface: Tool LayerPropagation: Single Hop

A server-side request forgery flaw exists in the mcp-wiki/wiki-summary component of AIAnytime Awesome-MCP-Server, where the 'url' argument passed to an MCP tool is not validated before the server fetches it. A remote attacker can supply this MCP-exposed tool with an internal or attacker-controlled URL to make the server issue requests on their behalf, potentially reaching internal network resources. The vendor has been notified but has not responded or patched the issue.

commentaryvulnerability-researchai-agentsdual-useoffensive-securitytrend-observationSurface: ModelPropagation: None

This is a brief opinion/commentary post observing that AI agents are becoming effective at finding software vulnerabilities at scale, referencing a tweet and a prior LinkedIn post about a coming 'AI Vulnerability Cataclysm.' It does not describe a specific vulnerability, exploit, attack technique, or affected system, so it does not constitute a genuine, actionable security threat in itself.

prompt-injectionadversarial-imagecross-model-attackmemory-toolindirect-injectionclaudechatgptmultimodalASI03 · Identity SpoofingAML.T0051AML.T0054Surface: MemoryPropagation: Single Hop

A researcher demonstrated that an image generated by ChatGPT could act as a carrier for an indirect prompt injection that hijacked Claude Opus 4.7's memory tool, causing it to persist false memories into future conversations. This shows that even hardened, reasoning-heavy models remain vulnerable to multimodal adversarial inputs crafted using puzzle-like framing to bypass safety reasoning.

TOCTOUcomputer-use-agentrace-conditionbrowser-agentChatGPT-OperatorUI-confirmation-bypassASI06 · Memory PoisoningSurface: PlannerPropagation: None

This research describes a time-of-check-to-time-of-use (TOCTOU) attack against computer-use AI agents like ChatGPT Operator, where a malicious page or element changes between the moment the agent evaluates it and the moment it acts, causing the agent (and a supervising human) to click or execute something different from what was reviewed. The author reproduced a previously disclosed Google-reported vulnerability and demonstrated it live at a security conference. This is a legitimate and impactful vulnerability class for autonomous browser/UI-driving agents.

informationalcoding-agentLLM-libraryno-vulnerabilityblog-postASI05 · Unsafe Code ExecutionSurface: Tool LayerPropagation: None

This is a Simon Willison blog post announcing an experimental alpha release of a Python coding agent library (llm-coding-agent) built on his LLM framework. It describes an AI-generated coding agent with file editing, shell execution, and file search tools, but the content is purely a release announcement with no evidence of a vulnerability, exploit, or malicious activity. The presence of powerful tools like execute_command and edit_file is inherent to any coding agent design and is explicitly disclosed by the author (including a --yolo flag), not a hidden threat.

claude-codeagent-memorysubagent-delegationmodel-routingnon-maliciousblog-postSurface: MemoryPropagation: None

This is a benign blog post by Simon Willison describing a legitimate workflow tip for Claude Code: instructing the agent to use its own judgement to delegate coding subtasks to cheaper/lower-power models via subagents, and how Claude persisted this preference as a memory file. There is no evidence of prompt injection, malicious payloads, or exploitation here; it simply illustrates that coding agents can write persistent memory files based on user instructions.

tool-usereliabilitycoding-agentsschema-mismatchclaudepinot-a-security-vulnerabilitySurface: Tool LayerPropagation: None

This report describes a reliability/compatibility quirk, not a security threat: newer Claude models (Opus 4.8, Sonnet 5) sometimes emit malformed tool call arguments with invented fields when used with third-party coding harnesses like Pi, likely due to RL training tuned specifically for Anthropic's own Claude Code edit tool. There is no malicious actor, injection, or exploitation involved—this is a model behavior/engineering problem causing failed tool calls and retries.

ai-pair-programmingcode-reviewsqlite-utilssoftware-qualitynon-securitySurface: Human InterfacePropagation: None

This is a blog post by Simon Willison describing how he used an AI coding assistant (Claude, referred to as 'Fable') to review and improve the sqlite-utils library ahead of a 4.0 stable release. The AI identified real software bugs, including a transaction-commit bug causing data loss, but this is a legitimate development workflow with no indication of prompt injection, tool poisoning, or any adversarial agent behavior.

red-teamingagentic-RAGmultimodalimage-injectiontext-poisoningorchestrator-manipulationMCTSresearchbenchmarkASI01 · Goal HijackingAML.T0051AML.T0054AML.T0043Surface: PlannerPropagation: Single Hop

This is an academic red-teaming paper (not an active exploit) introducing MIRROR, a search-based framework that automatically generates diverse, non-duplicated attacks against multimodal agentic RAG systems across text poisoning, image injection, direct-query, and orchestrator-manipulation surfaces. It demonstrates high attack success rates, notably 97% against orchestrator-level tool manipulation, highlighting that current agentic RAG defenses are weak across multiple input modalities and pipeline stages. Because it is a research disclosure with an accompanying benchmark rather than an in-the-wild campaign, it is rated medium severity as a forward-looking risk indicator rather than an active incident.