Shostack + Friends Blog

 

The Pentagon, China, AI and PHANTOM-B

Adam Shostack, Shostack + Associates

Could PHANTOM-B have stopped this close call? an ai image of an ai ship

CNN has an explosive story with US military having a close call after using AI for false intelligence report. There were planes in the air and the United States was ready to initiate military action against a Chinese ship:

It was only just before the planned operation that officials dug deeper into the report put together by a special operations command analyst and found it had been generated with the help of artificial intelligence (AI) — and that a chatbot the analyst had used inaccurately identified the material the ship was carrying. CNN was not able to learn what the misidentified cargo was.
The report, according to one of the sources, was “entirely false.” But it also “almost started a war,” the source said. Any US operation against a Chinese vessel could have risked spiraling into an armed conflict between the two nations.

Now, it’s unclear from CNN’s reporting if it was a Chinese-flagged merchant ship, or a Chinese military vessel. (The other articles I’ve seen report on what CNN has reported, and have not developed new facts; I expect more will emerge.) While either would be bad, attacking a military vessel would be more likely to lead to a broader conflict. That broader conflict is yet more likely while the war with Iran has left the US military with depleted inventories of missiles and other supplies.

Whatever sort of ship it was, the issue here is how the military is using AI, and here, I want to talk about PHANTOM-B and threat modeling LLM-centered applications. (You can get a broad overview from our PHANTOM-B blog post or from our whitepapers page.) The “O” in PHANTOM-B stands for Overreliance, and it means just that: relying on the model. Apparently, the command analyst had generated a report, and not checked the key facts.

In the August paper, I wrote “If you let your model produce results without oversight, you’re going to be at least embarrassed, if not worse.” I wasn’t thinking that the Pentagon’s new AI strategy would go so far as letting AI make deployment or mission decisions without proper oversight, because after all, some commander is putting troops in harm’s way, and I would hope that commander would meet his duty to those soldiers. (That strategy is described in the CNN article linked at the top.)

That hope drives demand, and that demand leads to pressure to deploy.

Making consequential decisions is at the heart of “what can go wrong” with LLMs, and the list of ways over-reliance plays out goes on and on: This person gets arrested. We shouldn’t hire Asian candidates. We should lay off people who took maternity leave.

The reason we made PHANTOM-B a small, usable tool is that there’s tremendous hope that LLMs will let us make decisions faster and better. That hope drives demand, and that demand leads to pressure to deploy. If we have simple ways to analyze a plan, that simple way is more likely to be used.

PHANTOM-B provides a short set of human-level prompts to consider what can go wrong. (It’s literally designed to fit on a wallet card; you can get those PHANTOM-B wallet cards from our partner, Cybersecurity Games.) It’s free and creative-commons licensed.

That prompting to consider over-reliance can drive interventions or checkpoints in the design of a system. For example, a checklist item of “have you checked the key facts” or “will you stand by the recommendations in this report?” It could lead to the analyst or their commander asking “what here came from AI.” The crucial improvement is a recognition that over-reliance is a threat to the system.

Or, you know, you could start a war. Your choice. (Please don’t start more wars.)

Image: Overly reliant on Gemini.