Within the Rules: SGLang's SafeUnpickler Bypass Is the Fourth Critical AI-Infra CVE in 18 Days
CVE-2026-86793 lets an unauthenticated attacker run arbitrary code on SGLang inference servers by chaining two permitted builtins functions — no patch exists yet. It is the fourth critical CVE to hit the AI stack since August 25.
On September 11, 2026, VicOne published a technical analysis of CVE-2026-86793, a vulnerability in SGLang — one of the most widely deployed open-source frameworks for serving large language models — that allows a remote, unauthenticated attacker to execute arbitrary code on an affected server. What makes the bug notable is not just its severity. It is the elegance of the bypass: the attacker never breaks SGLang’s security rules. They simply stay within them.
A safeguard built to stop exactly this
SGLang has been here before. The framework’s SafeUnpickler class exists specifically to mitigate CVE-2025-10164, an earlier arbitrary-code-execution flaw that rode on unsafe pickle deserialization. Pickle is Python’s native serialization format, and it is a notorious attack surface in the machine-learning ecosystem: when a pickle byte stream is deserialized, the unpickler may import modules and resolve classes or functions to reconstruct the object graph. Control the byte stream, and you can make the target machine call a function of your choosing.
SafeUnpickler was SGLang’s answer. Its find_class method restricts which Python modules, classes, and functions can be resolved during deserialization, using two mechanisms: an allowlist of “safe” module prefixes (ALLOWED_MODULE_PREFIXES) and a denylist of specific module-and-name pairs that must be blocked (DENY_CLASSES).
The intent was sound. The implementation had two flaws that, combined, proved fatal.
Flaw one: a prefix that permits everything
The first problem is that the allowlist includes the prefix builtins. Because prefix matching accepts any name from the builtins module unless that exact name is explicitly denied, the effective policy became “all of Python’s built-in functions are allowed, minus a short blocklist.”
The second problem is that the blocklist is incomplete. It blocks the obvious guns — eval, exec, compile, open — but it does not block __import__ or getattr.
Those two omissions are all an attacker needs. The published proof of concept chains them in three moves:
builtins.__import__("os")— import the os module.__import__is a builtins function, so the allowlist waves it through.builtins.getattr(os_module, "system")— retrieveos.systemviagetattr. Again, a permitted builtins function.- Call the retrieved function with attacker-controlled arguments — in the PoC,
os.system("touch /tmp/poc_confirmed").
The critical detail is what find_class inspects. It sees only the module-and-name pairs passed to it during resolution: builtins.__import__ and builtins.getattr. It never sees ("os", "system") being resolved directly, because that resolution happens inside a legitimate function call, not through the deserializer’s class-lookup path. The denylist never triggers. As VicOne’s analysis puts it: the restrictions are bypassed not by breaking the rules, but by staying within them.
An admin endpoint with optional authentication
The gadget chain needs a delivery route, and SGLang provides one: the /update_weights_from_tensor endpoint, marked AuthLevel.ADMIN_OPTIONAL. This endpoint is designed to accept serialized tensor data — a base64-encoded pickle payload in the request — for in-place model weight updates. When no API key or admin API key is configured, it accepts requests without any authentication.
An attacker who can simply reach the HTTP server can POST a crafted payload and land code execution. No credentials, no sandbox, no user interaction.
Two months from report to disclosure — and still no patch
The disclosure timeline is its own story. VicOne researcher Reuel Magistrado submitted a private GitHub Security Advisory to the SGLang maintainers and followed up by email. A maintainer acknowledged the report via Slack on July 2, 2026 — but provided no patch and no remediation timeline. On July 16, with the project unresponsive, the issue was escalated to CERT/CC for coordinated disclosure. CERT/CC validated the vulnerability and assigned the CVE on September 8. VicOne published its full technical analysis on September 11.
As of publication, no official SGLang patch exists. VicOne’s recommended mitigations, developed by Magistrado but not yet validated as an official fix, are structural: abandon broad prefix matching for the builtins module in favor of an explicit class allowlist containing only the exact classes required for tensor serialization — or drop the generic allowlist approach entirely. Until a patch ships, organizations running SGLang should configure API-key authentication on the server and restrict access to the endpoint to trusted networks.
The fourth critical CVE in 18 days
CVE-2026-86793 does not stand alone. Forkast’s analysis frames it as the fourth critical data point in an 18-day window:
- August 25 — an Ollama inference backend flaw (CVSS 8.1): a backend bound to all network interfaces with Host-header validation disabled could be fully hijacked via DNS rebinding, triggered by a single visit to an attacker-controlled webpage.
- September 8 — DeepSeek Harness (CVSS 9.4): the first confirmed agent-runtime sandbox escape. An unauthenticated local API plus an OS sandbox that left loopback networking open meant one
curlfrom inside the container escalated the agent to “danger-full-access” with approval prompts disabled. - September 8 — IBM Langflow (CVSS 9.8): unauthenticated RCE during graph construction, via an incomplete denylist that omitted process-spawning primitives and a code parser that passed return-type annotations straight to
eval. - September 11 — SGLang (CVE-2026-86793): the SafeUnpickler bypass described above.
The throughline is uncomfortable. Each of these frameworks — NemoClaw, DeepSeek Harness, IBM Langflow, SGLang — was built to make high-velocity interaction between models, code, and local environments easy, and each prioritized ease of integration over strict access control. DeepSeek Harness alone racked up 215,000 GitHub stars within weeks of release, meaning insecure defaults were being deployed at scale before security teams could put compensating controls in place.
And the target layer keeps moving deeper. The authentication gap has migrated from middleware through enterprise software, VPN infrastructure, and network management planes, into agent runtimes — and now into the inference server itself: the component that loads model weights, processes tensors, and serves predictions. When authentication fails at this layer, the blast radius includes every model the server hosts and every application downstream of it.
The disclosure rate tells the same story quantitatively: from roughly one critical AI-infrastructure CVE per month in 2025 to roughly one per week by the third quarter of 2026. For teams operating LLM infrastructure, VicOne’s closing principle is the operational takeaway: secure deserialization cannot rely on blocking known-dangerous functions. Narrowly scoped allowlists for serialized data, and mandatory authentication on administrative endpoints, are not optional hygiene anymore — they are the difference between an inference server and someone else’s compute.
Sources
- [1] https://vicone.com/blog/cve-2026-86793-sglang-bypass-could-let-attackers-run-code-on-ai-servers/
- [2] https://forkast.news/four-weeks-four-critical-cves-ai-inference-infrastructure-is-now-a-regular-target/
- [3] https://www.cve.org/CVERecord?id=CVE-2026-86793
- [4] https://cvj.ai/briefing/crypto-news/ai-stack-under-siege-four-critical-cves-in-18-days-expose-inference-server-flaws/