Four Models Vote, No Humans Admitted: Inside CLOSEDQUORUM, the First Autonomous AI C2 Implant
Cisco Talos documents CLOSEDQUORUM, a Windows implant whose command-and-control is a quorum of four commercial LLMs — DeepSeek, Qwen, Mistral and Gemini — voting on each attack step with no operator in the loop.
For years, the reassuring line about AI-assisted cyberattacks was that the human stayed in the loop. AI could write better phishing lures and generate more malware variants, but a person still picked the targets, ran the tooling, and decided what happened next. On September 22, 2026, Cisco Talos published an analysis that quietly retires that comfort: a Windows implant called CLOSEDQUORUM that delegates its command-and-control decisions to a panel of four commercial large language models — and executes whatever they agree on. To Talos’s knowledge, it is the first publicly documented Windows implant to use LLMs as autonomous, tactical C2.
What CLOSEDQUORUM is
CLOSEDQUORUM is a 16.4 MB, 64-bit Windows executable compiled in Go (with CGO enabled, letting it mix in C code for direct Windows system calls). None of its offensive capabilities are novel on their own: LSASS memory dumping for Windows credentials, theft of saved browser passwords from Chrome, Edge and Firefox, extraction of crypto wallet data from MetaMask, Exodus and Ethereum wallet paths, process injection via Early Bird APC or process hollowing, and persistence through Registry Run keys, scheduled tasks and WMI event subscriptions.
What is new is the architecture. Traditional malware needs an attacker-operated C2 server: a domain, an IP, a protocol, a listener. That infrastructure is attributable, blockable, expensive to rotate, and trackable through certificate transparency logs and threat intelligence feeds. CLOSEDQUORUM has none of it. Instead, its C2 infrastructure is the commercial AI API ecosystem. On each decision cycle, the binary queries up to four LLM providers — DeepSeek, Qwen, Mistral and Google Gemini — endpoints used legitimately by thousands of applications every day.
How the quorum works
The name is precise. A quorum is a decision-making body requiring a minimum number of participants to act; CLOSEDQUORUM’s quorum is up to four LLM providers, and the session is closed — no humans are admitted. Internally, a component Talos calls the ModelOrchestrator queries the four providers one by one. Each response is deserialized into a typed Go struct and collected into a slice of decisions, which an interModelDiscussion() function resolves by plurality vote: each provider’s decision increments a counter map, and the highest count wins.
The multi-provider design is deliberately resilient. Individual refusals, timeouts, guardrail triggers and malformed responses lose their veto power when three other models still return a verdict. If every model fails, the fallback “decision” is a consensus string with no capability handler attached, so the loop just sleeps and retries rather than doing something reckless.
The models are not free-form chatting, either. They are constrained to a typed JSON schema — a small attack-decision language. The system prompt extracted from the binary reads: “You are an advanced malware strategist. Provide ONLY executable decisions.” The prompt template enumerates the choices and substitutes target context at runtime: hostname, OS architecture, CPU count, Windows version, admin status, plus a refreshed target process. The valid decisions map directly onto capability handlers: steal fires LSASS dumping, browser credential theft and wallet extraction simultaneously; inject generates shellcode and then picks process hollowing or APC injection based on the model-supplied exploit type; persist establishes a foothold; move has no handler in the distribution build. Anything that doesn’t parse as a known decision is discarded.
Tie-breaking is fully deterministic and biased toward DeepSeek: the max-finding loop iterates in submission order with a strict comparison, so in any tie the first-queried provider’s vote wins — DeepSeek, then Qwen, then Mistral, then Gemini.
A credentials-as-a-service business model
A crucial caveat: Talos has no confirmation of in-the-wild deployment. The publicly distributed binary is an inert template — API keys initialize to dummy_api_key and the Discord webhook to dummy_webhook_url. But development builds demonstrate build-time injection of real provider credentials, and artifacts from the binary connect its developer to carding-related postings on criminal forums dating back to 2025. The inferred operating model is a定制 service: the developer compiles a custom executable for each operator with their webhook and API keys baked in; the operator handles delivery; stolen data lands in the operator’s Discord channel, AES-256-GCM encrypted with a date-derived key, split into 1,900-byte Base64 segments posted at one-second intervals.
The encryption detail is telling about the criminal market’s maturity — or lack of it. Because the symmetric key derives from the current date, the developer’s infrastructure can decrypt any operator’s exfiltrated haul whenever it knows the date, which it always does. It is obfuscation, not a genuine confidentiality separation between developer and operator.
Effort displacement: the third dimension
Talos frames the finding with a concept worth remembering: effort displacement. AI’s offensive impact so far has run along two dimensions — speed (faster lure generation) and scale (more code variants). Both leave the human operator in place. Effort displacement is the third dimension: transferring an entire phase of the attack from the human to the system. Human operators are bounded by attention, working hours and cognitive load. An implant that can execute a phase of the attack chain keeps running when the attacker logs off and sleeps. It compounds speed and scale, because the human-in-the-loop is no longer the bottleneck.
That framing also explains why defenders should not panic. Autonomy trades human limitations for model and infrastructure limitations: provider refusals, rate limits, malformed output, a deterministic tie-break order, a constrained action schema, and total dependence on commercial APIs. Each is a failure mode and a defensive opportunity.
How to detect it
Talos’s guidance is behavioral, not indicator-based. No single IOC identifies the architecture, but the combination is distinct: AI-provider API traffic originating from an unexpected Windows executable; requests to several model providers within a short interval; structured prompts containing host context or offensive capability language (visible only via TLS inspection or provider-side telemetry); classic techniques like LSASS access, process injection or WMI persistence; Discord webhook communication from the same host; and execution repeating at randomized 5–15-minute intervals. Legitimate apps may contact DeepSeek or Discord independently. Far fewer should contact several AI providers while also touching LSASS and creating WMI subscriptions.
Alongside the disclosure, Talos open-sourced CAIRN (Cognitive Artifact Intelligence Research Network), the research toolkit that surfaced CLOSEDQUORUM, designed to hunt, classify and track AI-integrated malware. Talos describes CLOSEDQUORUM as the first in a series of findings — the threat class will span from proof-of-concept experiments to active campaigns.
Why it matters
The honest read is that CLOSEDQUORUM is not sophisticated malware; Talos says as much. Its significance is as an existence proof: removing the operator from a bounded phase of an intrusion is achievable today, with currently available models and ordinary paid API access. The architectural trick — encoding tactical attack logic as model-readable context and converting structured model output directly into execution — is scaffolding that can be transplanted to other adversary objectives. Talos’s closing point is the right one: the autonomy arc is only beginning, and defenders currently have an open window to build the detections and response strategies before autonomous operations become more capable and more widespread. Windows into this transition, like this one, are exactly how that window gets used.
Sources
- [1] https://blog.talosintelligence.com/the-closed-quorum-inside-the-first-reported-autonomous-ai-c2-implant/
- [2] https://www.bleepingcomputer.com/news/security/new-closedquorum-windows-malware-uses-ai-for-attack-decisions/
- [3] https://www.theregister.com/security/2026/09/22/windows-closedquorum-malware-uses-ai-models-to-autonomously-select-post-compromise-actions/5298435
- [4] https://www.helpnetsecurity.com/2026/09/22/cairn-open-source-framework-ai-malware-closedquorum/
- [5] https://www.scworld.com/news/1st-autonomous-ai-c2-implant-uses-panel-of-models-to-vote-on-next-task