ravi@rajput:~$

Threat modeling agentic AI orchestration: a component-by-component map

TL;DR

Model the orchestration stack as nodes and edges, then ask which edges cross a trust boundary. Six do: untrusted content into context, planner into tool invocation, agent identity into tool authorization, memory write into memory read, agent into peer agent, and decision into execution. Every significant agentic attack is an abuse of one of those six. The model is not where the vulnerability lives.

Draw your agentic system as nodes and edges. Then mark every edge that crosses a trust boundary.

Most teams have never done this, which is why most agentic threat models are a list of prompt-injection payloads. The payload is not the vulnerability. The edge it travels is.

The reference architecture

agentic-orchestration — trust boundaries
UNTRUSTED CONTENT user message retrieved document tool response peer agent message tool definition all of it is instructions to the model B1 planner / reasoner goal decomposition memory cross-session state RAG / vector store retrieval corpus B4 B2 tool / function layer invocation + schema agent identity entitlements B3 MCP / tool servers third-party surface peer agents A2A handoff B5 PRIVILEGED database payments email / send admin API human gate irreversible only B6 TRUST BOUNDARIES B1 untrusted content → model context B2 planner output → tool invocation B3 agent identity → tool authorization B4 memory write → memory read B5 agent → peer agent B6 decision → execution
Six trust boundaries. Red dashed edges cross one. Every significant agentic attack is an abuse of B1 through B6 — none of them are the model itself.

Six boundaries. That is the whole threat model, and each one has a characteristic failure.

Boundary by boundary

B1 — untrusted content into model context. Everything in the left-hand zone arrives on the same channel as your system prompt: user messages, retrieved documents, tool responses, peer agent messages, and the tool definitions themselves. The model has no structural marker separating policy from data, so all of it is instructions. The attacker rarely speaks to your model directly; they leave text where it will read.

B2 — planner output into tool invocation. The planner emits a structured call. If the tool layer executes it because it is well-formed, the planner’s reasoning has become your authorization decision. Planner output is untrusted input to the tool layer, and a policy enforcement point belongs on this edge.

B3 — agent identity into tool authorization. This is the one that matters most and gets the least attention.

B4 — memory write into memory read. Memory turns a one-shot injection into persistence. Content written in one session is read as trusted context in the next, which is why memory poisoning is T1 in OWASP’s taxonomy rather than a footnote under injection.

B5 — agent into peer agent. Authenticating the sender is not authorizing the request. More on this below.

B6 — decision into execution. The gap where a human approval either exists or does not.

B3: where excessive agency actually lives

The agent holds tools and privileges the requesting user cannot reach directly. That is the point of the agent — it is more capable than the person asking.

Which means compromising the agent is a privilege escalation path by construction, and you built it deliberately.

confused-deputy — B3
user read own orders asks agent refund_order() read_all_orders() acts as itself payments API trusts the agent gap = everything the agent can do that the user cannot B3 — AUTHORIZATION GAP FIX scope agent entitlements to least privilege AND validate the prompting user is authorized for this specific action
The confused deputy. Scoping the agent narrowly is necessary and insufficient. If the agent can refund any order and any user can ask it to, you have built a confused deputy with good intentions.

OWASP’s mitigation is two-part and most implementations ship only the first half: least-privilege scoping of agent entitlements, plus per-request validation that the prompt-submitting user is authorized for the action being requested.

B5: lateral movement with no credential theft

This is the attack path the agent layer uniquely creates.

lateral-movement — B5
triage agent compromised reads untrusted email trusted peer enrichment agent queries CMDB trusted peer response agent isolates hosts EDR admin API privileged NOTHING WAS STOLEN no credential theft · no anomaly in identity logs · the trust was already provisioned blast radius = transitive closure of the trust graph
Lateral movement across agents. An attacker controlling one agent inherits its peer relationships and moves through them without stealing a credential — so there is nothing anomalous for identity monitoring to catch.

Two architectural consequences:

Authentication is not authorization. The receiving service must independently enforce the sender’s permissions rather than honouring a call because it arrived from a known peer. Message signing gives integrity only — it proves the message was not tampered with, not that the sender was allowed to ask. Transport encryption remains separately necessary.

Draw the trust graph. The blast radius of one compromised agent is its transitive closure. Most teams have never drawn it, and the exercise usually reveals that the agent reading untrusted content has a two-hop path to something privileged.

The component-to-technique table

ATLAS now has an agentic technique family mapping close to one-to-one onto these components. Verified against MITRE’s machine-readable dataset (ATLAS-2026.09.yaml):

ComponentBoundaryAttackATLAS
Tool / function layerB2Unauthorised invocationAML.T0053
MCP / tool serversB2Tool poisoning — definition, implementation, runtime responseAML.T0110 (.000 .001 .002)
Tool supply chainB2Poisoned third-party toolAML.T0010.005, AML.T0115.002
MemoryB4Context and memory poisoningAML.T0080 (.000)
RAG / vector storeB1Corpus poisoningAML.T0070
RAG / vector storeB1False entry injectionAML.T0071
RAG / vector storeB1Credential harvesting from corpusAML.T0082
Agent identityB3Config modificationAML.T0081
Agent identityB3Credentials from agent configAML.T0083
Tool definitionsB2Definition tamperingAML.T0084.001
Any tool edgeB2Exfiltration via tool invocationAML.T0086

AML.T0110 names Model Context Protocol connections explicitly — relevant if you are standing up MCP servers and wondering whether the risk is still theoretical.

RAG earns three distinct techniques because retrieval is now a first-class attack surface rather than a sub-case of data poisoning. Credential harvesting is the one that surprises people: a corpus accumulates whatever was indexed, indexing rarely audits for secrets, and the agent is an extremely effective search interface over exactly that material.

B6: the control that survives everything upstream

For transfers, administrative operations and anything irreversible, separate decision from execution architecturally, with approval bound to the exact actor, tool call and normalized parameters, and made single-use.

Each clause carries weight. Separating decision from execution means the component that decides cannot act. Binding to normalized parameters means approving refund order 4417 for £82.50, not refund an order — a generic approval is a blank cheque the model fills in later. Single-use prevents replay against a second call.

This is ordinary authorization engineering. It is also the only control on the diagram that holds when every control to its left has failed.

Which framework for which job

Four bodies of work, dividing the labour rather than competing:

  • OWASP Agentic Security Initiative — Agentic AI: Threats and Mitigations, v1.1 (Dec 2025), seventeen threats T1 Memory Poisoning through T17 Supply Chain Compromise, each paired with named mitigations. Use it for threat-to-control rows.
  • MITRE ATLAS — technique IDs. Use it so agentic threats enter the same detection-engineering pipeline as everything else instead of living in a document nobody queries.
  • CSA MAESTRO — seven layers plus distinct threat profiles per orchestration pattern. Single-agent, multi-agent, hierarchical and distributed-ecosystem deployments differ, and most writing collapses them.
  • OWASP LLM Top 10 — shared risk vocabulary.

STRIDE, PASTA, LINDDUN and the rest still cover the conventional surface competently. They do not decompose a system in a way that surfaces B4 or B5.

Start here

Draw the node-edge diagram for your own system. Mark B1 through B6. Attach the ATLAS IDs. Then walk OWASP’s T1–T17 against it and mark which threats your architecture actually admits.

Two findings are near-universal in that exercise: agent identity is a single shared service account, and nobody has drawn the trust graph.

Both are fixable before they are incidents. Neither is an AI problem.


Citation notes. OWASP shipped a 2026 LLM Top 10 in August 2026 reordering categories — Excessive Agency to #3, Vector and Embedding Weaknesses to #9, System Prompt Leakage renamed Hidden Context Exposure — so LLMxx:2025 numbering now reads stale. The ASI document is v1.1 with seventeen threats, not the widely quoted fifteen. AML.T0104 no longer exists in ATLAS v6; use AML.T0115.002. ATLAS deep links at atlas.mitre.org/techniques/<ID> return 404 to non-browsers because the site is a JS SPA — cite the atlas-data YAML for durability. ATLAS itself never uses the terms "excessive agency" or "confused deputy"; those are OWASP labels, and the mapping between them here is mine — no primary-source crosswalk between ATLAS, OWASP T1–T17 and MAESTRO layers currently exists. MAESTRO is a named-author CSA article rather than a ratified standard. Domain-specific threat variation (healthcare, banking, ecommerce, platforms) is deliberately out of scope here and is the subject of separate work.

Agentic AIThreat ModelingMITRE ATLASOWASPMAESTROAI Security

Frequently asked

Which threat modeling framework should I use for agentic AI?
They are complementary. OWASP's Agentic Security Initiative taxonomy gives seventeen threats paired with mitigations, MITRE ATLAS gives machine-readable technique IDs that integrate with existing detection engineering, and CSA's MAESTRO gives a seven-layer architectural canvas with distinct threat profiles per orchestration pattern. Classic frameworks like STRIDE still cover the conventional surface but do not decompose a system in a way that surfaces memory poisoning or inter-agent trust abuse.
Where does excessive agency actually live in an agentic architecture?
At the boundary between agent identity and tool authorization. The agent holds tools and privileges the requesting user cannot reach directly, so compromising the agent is a privilege escalation path by construction. The mitigation is least-privilege scoping of agent entitlements plus per-request validation that the user who submitted the prompt is authorized for the action being requested.
Is authenticating an agent enough to authorize its requests?
No. Authenticating the sending agent is explicitly insufficient for authorization in multi-agent orchestration. The receiving service must independently enforce the sender's permissions rather than trusting a call because it arrived from a known peer. Message signing provides integrity only, so transport encryption remains separately required.
How does an attacker move laterally between AI agents?
By abusing the compromised agent's pre-existing trust relationships with peer agents. No credential theft occurs, so there is nothing anomalous in identity logs — the trust was already provisioned and is being used exactly as designed. The blast radius of one compromised agent is the transitive closure of its trust graph.
What is the single most effective control for high-impact agent actions?
Architecturally separating decision from execution, with approval bound to the exact actor, tool call and normalized parameters, and made single-use to prevent replay. Approving 'refund order 4417 for £82.50' is a control; approving 'refund an order' is a blank cheque the model fills in later.
← all posts