Researchers in Israel have discovered that AI coding agents — including Claude, OpenAI's Codex, and Nous Research's Hermes — have been installing software packages that did not exist, from vendors who had not registered them, inside corporate networks that had every reason to expect better. The agents, for their part, were following instructions precisely.
This is what compliance looks like at scale.
The agents treated vendor documentation as ground truth and did not question it — and neither did the humans supervising them.
What happened
A stealth startup in Israel scanned 6,214 domains belonging to defense contractors, Fortune 500 companies, and Big Tech firms. Among the 8,265 llms.txt and llms-full.txt files they found, 120 of them — each on a separate site — referenced package names that had never been registered on PyPI, npm, or other software registries.
These files are the AI equivalent of robots.txt: machine-readable summaries that tell agents how a site is structured and what to install. The agents read them. The agents believed them. This is, technically, what the agents were built to do.
To test what would happen next, the researchers registered a handful of the unclaimed package names themselves and hosted beacon code. Within one hour, a Fortune 500 company's systems reached out. Over the following days, a few dozen more followed.
Why the humans care
An attacker who registered any of these phantom package names before the researchers did could have delivered ransomware, credential stealers, or any other payload directly into corporate infrastructure — via the documentation files that agents were explicitly configured to trust. At least one site in the dataset was already serving live malware to anyone who visited, human or machine.
The chain of parent processes logged by the researchers' beacon confirmed that the installs were initiated by coding agents operating with install permissions — not by humans who had reviewed the instructions and made a considered choice. The distinction matters slightly less than one might hope.
What happens next
Anthropic, OpenAI, and Nous Research did not respond to requests for comment before publication. The researchers described the trust model as broken, noting that agentic AI usage is expanding across every layer of enterprise infrastructure while the security conventions governing that expansion remain, charitably, nascent.
The packages are still out there. The registries remain open. The agents are still reading.