Notes
Field notes on agentic AI security, LLM red teaming, ML supply-chain security and offensive research — written by a practitioner who builds these systems and attacks them.
-
An agentic AI red team checklist: 222 tests, WSTG-style
A forced-order, 222-test checklist for authorized agentic AI engagements — 20 categories from recon to voice, each test mapped to OWASP ASI, the LLM Top 10, the MCP Top 10 and a MITRE ATLAS tactic, with how-to-test, tools and a detection mode. Free download, built to work top-to-bottom like the OWASP WSTG.
-
Excessive agency is an authorisation bug wearing an AI costume
The severity of an LLM compromise is set by what the agent is permitted to do, not by how it was tricked. Scoping tools to least privilege converts most agent vulnerabilities from critical to cosmetic — and it is ordinary authorisation work, not AI work.
-
LLM cost is a security control, not a finance problem
Running every security alert through your most capable model is a denial-of-wallet vulnerability and a detection gap at the same time. Severity-tiered routing cut our per-alert cost 30× — and the reason it works is a security argument, not a budget one.
-
Prompt injection is not a prompt problem
Prompt injection cannot be fixed with better prompts because LLMs have no channel separation between instructions and data. The durable mitigations are architectural: least-privilege tool scopes, output handling, and human gates on irreversible actions.
-
The agentic AI regulatory map is mostly blank — and one set of supervisors said so
Two regulators have published an agentic-specific risk taxonomy: FINRA and the FSB. US banking supervisors expressly carved agentic AI out of scope in April 2026. For ecommerce, platforms and the cross-cutting instruments, primary-source evidence does not yet exist.
-
Threat modeling agentic AI orchestration: a component-by-component map
An agentic system has six trust boundaries and most teams have drawn none of them. This maps the orchestration stack node by node — planner, tool layer, memory, RAG, inter-agent messaging, agent identity — to the attacks each edge carries and the MITRE ATLAS techniques that name them.
-
Your model file is a code execution primitive
Downloading a model from a public hub is running untrusted code. Pickle deserialisation, Keras Lambda layers (CVE-2024-3660) and GGUF parsing all give an attacker execution before inference ever starts — and most ML pipelines scan none of it.