← All posts / Research

227 Dangling Install Commands: The llms.txt Files That Turned Corporate Docs Into an Attack Surface

Researchers scanned 6,214 corporate domains and found AI-facing llms.txt files pointing at 227 unregistered packages and domains — and proved Claude, Codex, and Hermes agents inside Fortune 500 firms would execute them.

227 Dangling Install Commands: The llms.txt Files That Turned Corporate Docs Into an Attack Surface

On August 27, 2026, Ars Technica senior security editor Dan Goodin published a story with an unusual cast of villains: nobody. The threat it described wasn’t a hacker group or a nation-state toolkit — it was 227 install commands sitting in corporate documentation files, pointing at software packages and domains that no one owned. Anybody could claim them. And inside some of the world’s largest companies, AI coding agents were already running those commands.

The research, conducted by a stealth-stage team in Israel including researcher Alon Hertz, scanned 6,214 live domains belonging to defense contractors, Fortune 500 companies, and Big Tech. The target wasn’t application code or network infrastructure — it was llms.txt and llms-full.txt files: an emerging convention, modeled on robots.txt, where websites publish machine-readable summaries and setup instructions specifically for AI agents to consume.

What they found reframes a classic supply-chain weakness for the agentic era. And their proof of concept — which triggered callbacks from real Fortune 500 networks within an hour — shows the problem isn’t theoretical. It’s already live, with at least one in-the-wild malware package documented.

What the researchers found

Of the 8,265 llms.txt and llms-full.txt files discovered across the scanned domains, 120 files — each on a different site — referenced at least one unregistered package or unclaimed domain. In total, the misconfigured files contained 227 commands instructing readers to install packages that don’t exist on PyPI or npm, or to visit domains nobody has registered.

Typical entries looked innocuous: “Installation: pip install [redacted]” or “npm install [redacted].” One file referenced a test framework hosted on a domain that was available for anyone to buy. The entries are traps not because they’re malicious, but because they’re dangling — orphaned references to namespaces their owners never claimed.

The origin of these broken entries varies. Some predate the AI era entirely, having been copied from older non-LLM documentation files and written by humans. Others were likely generated by AI tools themselves — hallucinated package names, or content scraped from untrusted sources that the generating model couldn’t distinguish from legitimate instructions. In both cases the outcome is identical: an authoritative-looking file, served over HTTPS from a company’s official domain, telling agents to fetch code from an unclaimed slot.

The proof of concept: code execution inside Fortune 500s

Finding dangling references is one thing. Proving that real agents would act on them is another. The researchers registered a handful of the unclaimed names and hosted packages that, when executed, phoned home to the researchers’ server.

Within an hour, they received their first callback — from a Fortune 500 company. Over the following days, a few dozen more arrived, including several more Fortune 500 networks and numerous startups. The beacon also captured the chain of parent processes behind each install, revealing the executors: coding agents including Anthropic’s Claude, OpenAI’s Codex, and Nous Research’s Hermes, running inside those corporate networks.

Anthropic, OpenAI, and Nous Research did not respond to Ars Technica’s requests for comment by publication time.

The researchers’ diagnosis is blunt: “The trust model is broken. Agents treat vendor docs as ground truth and don’t question them — and neither do the humans supervising them. Agentic AI usage is exploding, and agents are spreading across every layer — SaaS, cloud, endpoint. As they multiply, so does the supply-chain surface, and today’s guards don’t cover it.”

The live attack: malware via npx on clerk.com

The scariest part of the findings isn’t the PoC — it’s that at least one real-world attack has already used the pattern.

The researchers found an LLM file hosted on the legitimate website of Clerk, a widely used authentication vendor, containing the command npx clerk-next-fix-auth-protection. The npx tool is particularly dangerous here: unlike a conventional install, it fetches a package into npm’s cache and executes its exposed binary immediately, without adding anything to the project’s dependency manifest — meaning the action leaves a fainter audit trail.

Someone had already claimed the once-empty package slot and used it to host live malware. Clerk has since resolved the problem, noting that agents which had already installed the legitimate @clerk/eslint-plugin binary weren’t at risk from the lookalike. It remains unclear whether the malicious package actually infected anyone.

Why every layer of trust fails at once

What makes this attack class genuinely novel is that it defeats defenses not by exploiting a bug, but by weaponizing correct behavior. The researchers walk through why no security layer fires:

  • To the agent, the file is served over HTTPS, on the company’s official domain, in a standardized AI-facing format the company published itself. “The file is the authority — that’s its entire purpose,” they write. When the file says pip install internal-tool, the agent doesn’t pause to check whether internal-tool belongs to the company, verify the namespace on PyPI, or notice the documentation link points to a domain that expired three months ago. “It just does what the file says.”
  • To the endpoint, it looks like a developer running a legitimate package manager — pip install from pypi.org, a domain every corporate proxy already allows, with a coding agent the company deliberately installed as the parent process. “No anomaly. No alert. The failure happens upstream, in the gap between the instruction and the execution.”
  • The trust chain is transitive. The llms.txt file doesn’t need to sit on the Fortune 500’s own website. If an agent trusts a partner’s docs, a vendor’s SDK reference, or a community setup guide, and that third party’s file points to an unclaimed package, the chain works identically.

A new name for an old boundary collapsing

Security researchers have spent two years documenting prompt injection — attackers deliberately planting malicious instructions in content an AI will read. This research shows the threat is structurally broader than adversarial injection.

“In a prompt injection, someone deliberately plants malicious instructions,” Hertz told Ars. “Here, the instruction itself can be completely benign and come from a legitimate source — a real company’s own documentation — with no malicious actor involved at the time it’s written. The danger comes later, when the package or domain it points to is abandoned and someone else claims it.”

As the researchers put it: “An agent doesn’t distinguish between a page and a command. Everything it reads is input, and every input is a potential instruction. Which means the entire corpus of published data that agents are now wired to consume has silently become an execution surface — and almost none of it carries the integrity guarantees we apply to actual code.”

The line between data and code — the boundary compilers, sandboxes, and code-signing regimes have policed for decades — doesn’t exist inside an LLM. “The Clerk case is the cleanest proof of it,” they write. “The command looked exactly like something the vendor would ship — because it was in the vendor’s own instruction file. The only thing missing was the name in the registry. Every layer of trust was intact except the one nobody thought to check.”

What to do about it

The findings carry direct action items for anyone operating agents with shell access:

  1. Verify namespaces before execution. Agents (or their harnesses) should confirm that a package referenced in third-party documentation actually belongs to the publishing organization before installing it — registry ownership, download history, publisher identity.
  2. Audit your own llms.txt. If your company publishes AI-facing instruction files, every install command and domain reference inside them is a promise about supply-chain ownership. Treat them like production code, not marketing copy.
  3. Don’t let agents run npx/pip install from untrusted docs unattended. The humans supervising these agents, as Hertz notes, tend to extend the same unquestioning trust the agents do.
  4. Treat registry monitoring as defensive infrastructure. Watching for registrations of names that appear in your public documentation converts this attack from free to expensive.

The deeper lesson is uncomfortable for an industry wiring agents into everything. Slopsquatting — attackers registering hallucinated package names — was already a known risk when models generate fake dependencies. This research shows the same attack works when models merely read documentation, at ecosystem scale, with the documents themselves published by the most trusted parties in the chain. The corpus agents consume became an execution surface silently. The guards that would make it safe mostly don’t exist yet — and until they do, the gap between instruction and execution is where the next generation of supply-chain attacks will live.