Threat modeling agentic AI orchestration: a component-by-component map
Model the orchestration stack as nodes and edges, then ask which edges cross a trust boundary. Six do: untrusted content into context, planner into tool invocation, agent identity into tool authorization, memory write into memory read, agent into peer agent, and decision into execution. Every significant agentic attack is an abuse of one of those six. The model is not where the vulnerability lives.
Draw your agentic system as nodes and edges. Then mark every edge that crosses a trust boundary.
Most teams have never done this, which is why most agentic threat models are a list of prompt-injection payloads. The payload is not the vulnerability. The edge it travels is.
The reference architecture
Six boundaries. That is the whole threat model, and each one has a characteristic failure.
Boundary by boundary
B1 — untrusted content into model context. Everything in the left-hand zone arrives on the same channel as your system prompt: user messages, retrieved documents, tool responses, peer agent messages, and the tool definitions themselves. The model has no structural marker separating policy from data, so all of it is instructions. The attacker rarely speaks to your model directly; they leave text where it will read.
B2 — planner output into tool invocation. The planner emits a structured call. If the tool layer executes it because it is well-formed, the planner’s reasoning has become your authorization decision. Planner output is untrusted input to the tool layer, and a policy enforcement point belongs on this edge.
B3 — agent identity into tool authorization. This is the one that matters most and gets the least attention.
B4 — memory write into memory read. Memory turns a one-shot injection into persistence. Content written in one session is read as trusted context in the next, which is why memory poisoning is T1 in OWASP’s taxonomy rather than a footnote under injection.
B5 — agent into peer agent. Authenticating the sender is not authorizing the request. More on this below.
B6 — decision into execution. The gap where a human approval either exists or does not.
B3: where excessive agency actually lives
The agent holds tools and privileges the requesting user cannot reach directly. That is the point of the agent — it is more capable than the person asking.
Which means compromising the agent is a privilege escalation path by construction, and you built it deliberately.
OWASP’s mitigation is two-part and most implementations ship only the first half: least-privilege scoping of agent entitlements, plus per-request validation that the prompt-submitting user is authorized for the action being requested.
B5: lateral movement with no credential theft
This is the attack path the agent layer uniquely creates.
Two architectural consequences:
Authentication is not authorization. The receiving service must independently enforce the sender’s permissions rather than honouring a call because it arrived from a known peer. Message signing gives integrity only — it proves the message was not tampered with, not that the sender was allowed to ask. Transport encryption remains separately necessary.
Draw the trust graph. The blast radius of one compromised agent is its transitive closure. Most teams have never drawn it, and the exercise usually reveals that the agent reading untrusted content has a two-hop path to something privileged.
The component-to-technique table
ATLAS now has an agentic technique family mapping close to one-to-one onto these components. Verified against MITRE’s machine-readable dataset (ATLAS-2026.09.yaml):
| Component | Boundary | Attack | ATLAS |
|---|---|---|---|
| Tool / function layer | B2 | Unauthorised invocation | AML.T0053 |
| MCP / tool servers | B2 | Tool poisoning — definition, implementation, runtime response | AML.T0110 (.000 .001 .002) |
| Tool supply chain | B2 | Poisoned third-party tool | AML.T0010.005, AML.T0115.002 |
| Memory | B4 | Context and memory poisoning | AML.T0080 (.000) |
| RAG / vector store | B1 | Corpus poisoning | AML.T0070 |
| RAG / vector store | B1 | False entry injection | AML.T0071 |
| RAG / vector store | B1 | Credential harvesting from corpus | AML.T0082 |
| Agent identity | B3 | Config modification | AML.T0081 |
| Agent identity | B3 | Credentials from agent config | AML.T0083 |
| Tool definitions | B2 | Definition tampering | AML.T0084.001 |
| Any tool edge | B2 | Exfiltration via tool invocation | AML.T0086 |
AML.T0110 names Model Context Protocol connections explicitly — relevant if you are standing up MCP servers and wondering whether the risk is still theoretical.
RAG earns three distinct techniques because retrieval is now a first-class attack surface rather than a sub-case of data poisoning. Credential harvesting is the one that surprises people: a corpus accumulates whatever was indexed, indexing rarely audits for secrets, and the agent is an extremely effective search interface over exactly that material.
B6: the control that survives everything upstream
For transfers, administrative operations and anything irreversible, separate decision from execution architecturally, with approval bound to the exact actor, tool call and normalized parameters, and made single-use.
Each clause carries weight. Separating decision from execution means the component that decides cannot act. Binding to normalized parameters means approving refund order 4417 for £82.50, not refund an order — a generic approval is a blank cheque the model fills in later. Single-use prevents replay against a second call.
This is ordinary authorization engineering. It is also the only control on the diagram that holds when every control to its left has failed.
Which framework for which job
Four bodies of work, dividing the labour rather than competing:
- OWASP Agentic Security Initiative — Agentic AI: Threats and Mitigations, v1.1 (Dec 2025), seventeen threats T1 Memory Poisoning through T17 Supply Chain Compromise, each paired with named mitigations. Use it for threat-to-control rows.
- MITRE ATLAS — technique IDs. Use it so agentic threats enter the same detection-engineering pipeline as everything else instead of living in a document nobody queries.
- CSA MAESTRO — seven layers plus distinct threat profiles per orchestration pattern. Single-agent, multi-agent, hierarchical and distributed-ecosystem deployments differ, and most writing collapses them.
- OWASP LLM Top 10 — shared risk vocabulary.
STRIDE, PASTA, LINDDUN and the rest still cover the conventional surface competently. They do not decompose a system in a way that surfaces B4 or B5.
Start here
Draw the node-edge diagram for your own system. Mark B1 through B6. Attach the ATLAS IDs. Then walk OWASP’s T1–T17 against it and mark which threats your architecture actually admits.
Two findings are near-universal in that exercise: agent identity is a single shared service account, and nobody has drawn the trust graph.
Both are fixable before they are incidents. Neither is an AI problem.
Citation notes. OWASP shipped a 2026 LLM Top 10 in August 2026 reordering categories — Excessive Agency to #3, Vector and Embedding Weaknesses to #9, System Prompt Leakage renamed Hidden Context Exposure — so LLMxx:2025 numbering now reads stale. The ASI document is v1.1 with seventeen threats, not the widely quoted fifteen. AML.T0104 no longer exists in ATLAS v6; use AML.T0115.002. ATLAS deep links at atlas.mitre.org/techniques/<ID> return 404 to non-browsers because the site is a JS SPA — cite the atlas-data YAML for durability. ATLAS itself never uses the terms "excessive agency" or "confused deputy"; those are OWASP labels, and the mapping between them here is mine — no primary-source crosswalk between ATLAS, OWASP T1–T17 and MAESTRO layers currently exists. MAESTRO is a named-author CSA article rather than a ratified standard. Domain-specific threat variation (healthcare, banking, ecommerce, platforms) is deliberately out of scope here and is the subject of separate work.