Threats and LLMs (Threat Model Thursday)
There are many ways to ask what can go wrong, and the useful skill is knowing which one matches what you're building.
Here, the series on the second edition of threat modeling takes a leap forward to Part III, which has four chapters on what can go wrong. Chapter 5, on STRIDE, is revised. Chapter 6 is largely new, and focuses on attack lifecycle models, mainly the Lockheed Kill Chain and ATT&CK, but also touches on ATLAS, FiGHT, a “Threat Matrix for Kubernetes”, SPARTA and VATT&CK. It also covers the relationship between chains and trees, and the widely varied language we use.
But the chapter that might be most exciting in this part is the one that led to PHANTOM-B, because the research that went into this chapter was pretty intense, and included reading probably a few thousand pages of LLM, ML and AI security, risk, threat, and governance documents and making sense of it in a way that’s helpful to readers.
The chapter kicks off with a simple model of “calling,” “running,” or “training” an LLM, because those are so influential on the threats you can do something about, and so that flavor of the question “what are we working on” permeates the chapter.
From there it covers a set of sets of what can go wrong with LLMs:
- PHANTOM-B
- MITRE ATLAS
- The Berryville Institute’s ARAs
- An image-specific taxonomy by Charlotte Bird
- Google’s SAIF
- NIST’s AIML (which is not their AIRMF)
- PROMISE TO MAP
Each is covered in depth. The chapter continues with a cornucopia of other ways to consider what can go wrong:
- Google’s empirical misuse list
- Meta’s Rule of Two for AI (and why it’s misnamed)
- Several approaches to AI red teaming
- Scholarly concerns including adversarial perturbation, gradient climbing and model theft
The chapter’s last technical element is a set of concerns that span outside of security: hallucination, bias or unfairness, and reliability, and especially the crucial relationship of security to reliability.
All of that does two things for you, the reader. First, it organizes the many approaches you might use, letting you select one that's appropriate to your situation. Second, each is presented as a scope: when it makes sense to consider it. For example, Bird’s taxonomy is useful for images. More generally, the question of “what can go wrong with an LLM” isn’t limited to prompt injection, and the different nature of the system may call for different threat modeling.
Image by midjourney: "A clean flat editorial illustration of a friendly, simple robot standing at a table, holding up a round glass lens to examine a single glowing orb. Three more lenses of different sizes float in front of the orb, each showing it a little differently. Simple shapes, soft gradients, subtle shadows, deep indigo purple background, lime green and bright orange accents, red-orange highlights. Curious, thoughtful mood, plenty of negative space. No text, no padlocks"