One Message, Root Shell: Unpatched CVSS 9.8 RCE in LMCache Exposes vLLM Inference Stacks
CVE-2026-105192 turns LMCache's multiprocess ZeroMQ port into an unauthenticated pickle-deserialization RCE that runs as root in official containers — and no fixed version exists.
The AI infrastructure stack has produced its first headline vulnerability of the season, and it is a nasty one. On October 7, 2026, JFrog’s security research team disclosed CVE-2026-105192, a critical remote-code-execution flaw in LMCache — the open-source distributed key-value cache layer that accelerates vLLM, one of the most widely deployed LLM inference engines. The bug scores 9.8 out of 10 on CVSS, and as of publication no fixed version exists. Every release from 0.3.9 (October 2025) through 0.5.5, the current stable, is affected — along with the 0.5.6 release candidates and the development branch itself.
For an industry that has spent 2026 racing to put AI agents in front of customers, the finding is a blunt reminder that the plumbing underneath those agents is often one misconfigured flag away from a full compromise.
How the bug works
LMCache speeds up LLM serving by caching KV-cache blocks so that repeated or shared prompt prefixes don’t have to be recomputed. In its multiprocess mode (also called distributed mode), the cache runs as a standalone server, and LLM worker processes reach it over the ZeroMQ messaging library. That design is exactly what multi-node deployments use to share a cache across machines.
The problem is what that server does with the messages it receives. The ZeroMQ ROUTER socket it opens — port 5555 by default — has no authentication whatsoever. Messages arriving on it are msgpack-encoded, but one extension code (code 1) is handed to DeviceIPCWrapper.Deserialize, which calls pickle.loads. Python’s pickle format can carry executable code, and deserializing it runs that code. Worse, the deserialize happens while the server is still decoding the request’s arguments, before the handler runs and before any check of the message type — so there is no validation layer the attacker has to get past.
The result: a single unauthenticated ZeroMQ DEALER message to the transport port executes code as the user the LMCache process runs as. And in the project’s official container images, that process runs as root. The flaw was found by Yuval Moravchick of JFrog.
Who is actually exposed
Whether a given deployment is remotely exploitable comes down to one setting. By default, the multiprocess server binds to localhost, which means another host cannot reach it. It becomes remotely reachable only when an operator starts it with a routable address via --host — which is precisely how multi-node deployments are supposed to let peers connect.
Here’s the uncomfortable part: LMCache’s own example Kubernetes deployment does exactly that, starting the server listening on every network interface. Operators who copied that manifest into production — a completely normal thing to do with a vendor’s reference config — are running a root-level, unauthenticated RCE on their inference cluster. A copy of LMCache running inside a single vLLM process doesn’t open the port at all.
JFrog’s guidance until a patch ships is operational rather than technical: don’t assign the multiprocess server a routable address; keep its port on the local machine or a trusted cluster network. A firewall that limits who can reach the port lowers risk but doesn’t eliminate it, since any host that can still open a connection can run code. Running the process under a non-root user and limiting container privileges blunt the blast radius, and enabling ZeroMQ authentication (or switching to a signed messaging protocol) closes the trust gap — but none of these is a substitute for a fix.
Adding to the opacity, LMCache has not published a security advisory for the flaw, and JFrog’s disclosure offers operators no way to determine whether a server has already been attacked.
More reports, and a fixed cousin
The disclosure didn’t land in a vacuum. On October 6 — the day before the CVE went public — a single GitHub account opened six additional security reports against LMCache. They allege unauthenticated access to cached data belonging to different tenants, plus several network services that execute commands without a login. Those reports remain unconfirmed proof-of-concept claims with no CVE and no fix, though one cited default has already changed: an admin HTTP server that listened on every interface in 0.5.5 now listens only on localhost in the 0.5.6 release candidates.
A related flaw in vLLM itself, by contrast, is already patched. Before version 0.30.0 (released September 22), a single request carrying a malformed cache_salt value could crash the engine on deployments using the LMCache multiprocess connector — a denial-of-service bug tracked as CVE-2026-105756, rated 6.5, with no code execution.
The pattern: ShadowMQ returns
The core mistake — handing data from an unauthenticated network socket to pickle — is the same class of flaw researchers cataloged across other AI inference frameworks in November 2025 under the name ShadowMQ. Whether LMCache’s code shares a common origin with those projects hasn’t been established, but the recurrence suggests something structural: as teams bolt high-performance IPC onto inference stacks, serialization formats that were tolerable inside a single trusted process are being exposed to networks that were never designed to be trusted.
The timing couldn’t be more pointed. This week also saw Pwn2Own Ireland pay out $388,500 for 32 zero-days — including full-takeover chains against OpenAI Codex and LiteLLM. LLM gateways, caching layers, and coding agents are no longer hypothetical attack surfaces; they are live ones, with bounties attached.
What operators should do today
If you run vLLM with LMCache in multiprocess mode: check your --host binding immediately. If the server is bound to a routable interface, treat the host as potentially compromised and move the port back to localhost or an isolated trusted network. Run the process as a non-root user with minimal container privileges, put the port behind a strict firewall allowlist, and watch the LMCache repository for a patched release — because until one ships, the only real fix is not being reachable.
The broader lesson for AI infrastructure teams is older than the industry’s current boom: deserialization of unauthenticated input is remote code execution. The fact that the input is a KV-cache block and the server is an LLM accelerator doesn’t change the math — it just changes who’s watching.