← All posts / Research

Fast Decisions, Slow Reasoning: Jev-Mem Splits Agentic Memory Into Two Systems and Cuts Query Latency 36.7%

A new paper swaps the LLM that usually steers agent memory for a small System-One decision model — 11% higher answer quality, 6.6× faster memory builds, and 0.93 s queries on LoCoMo.

Fast Decisions, Slow Reasoning: Jev-Mem Splits Agentic Memory Into Two Systems and Cuts Query Latency 36.7%

On September 21, 2026, a three-author team — Dongming Jiang, Yi Li, and Bingzhe Li — published Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents on arXiv, with the code already open on GitHub and a live demo on Hugging Face Spaces. Within hours it had become a top-trending story among AI practitioners. The reason is simple: the paper takes aim at a hidden tax that almost nobody budgets for — the cost of an AI agent remembering.

The problem: your memory system talks too much

Long-running agents — coding assistants that span months, personal assistants that span years — need persistent memory: preferences, events, connections across sessions. The awkward truth is that most modern memory architectures put a large autoregressive LLM in charge of the bookkeeping. Every time a memory is written, linked, or retrieved, the system asks a generative model to deliberate: How does this connect? Where should I search? Is this enough evidence?

That works, but it puts token generation — the single most expensive operation in the stack — on the critical path of operations that are fundamentally decisions, not prose. Memory builds take tens of minutes; queries take seconds; costs scale with the agent’s history. The Jev-Mem paper’s framing is blunt: these systems “rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expensive generation on the critical path of memory operations.”

The architecture: three planes, one division of labor

Jev-Mem’s answer borrows Daniel Kahneman’s old distinction between fast and slow thinking and makes it literal in the system stack:

  • System One — control plane. A small, non-generative decision model (Jev) handles every high-frequency judgment: memory typing, relation inference during construction, query routing across graph views, retrieval-budget allocation, candidate scoring, and — critically — adaptive stopping, deciding when enough evidence has been gathered. It doesn’t write text; it classifies and routes.
  • Memory plane — structured and multi-relational. Canonical observations keep their original text, timestamps, and provenance. Each memory participates in four overlapping graph views — semantic, temporal, causal, and entity — backed by vector and keyword indexes. The default profile retains every valid observation, so a detail whose importance was unclear at write time can still be found later.
  • System Two — reasoning plane. A conventional LLM (in the paper’s evaluation, GPT-4o-mini) is invoked only for complex reasoning and final answer synthesis.

The design principle worth stealing: the expensive model sees only the evidence; the cheap model runs the whole search. Retrieval becomes a loop — vector and keyword search supply anchors, the controller picks graph views, allocates traversal effort, scores candidates, and stops when evidence is sufficient or the expected value of more search drops below a threshold.

Two implementation details stand out for production-minded readers. First, Jev’s decisions are exposed as typed outputs — binary propositions (“Noul”) and categorical choices (“Choice”) — with ordinary code validating them and enforcing graph and retrieval limits. Second, every run emits inspectable traces: routing decisions, budgets, stopping reasons, cache hits, fallback events. For anyone who has debugged a memory system by reading generated text logs, that alone is a quality-of-life revolution.

The numbers

On the LoCoMo long-conversation benchmark, with GPT-4o-mini as the answer model and answer quality judged by LLM-as-a-Judge:

MethodOverall score ↑Build time (s) ↓Query latency (s) ↓
Full Context0.481N/A1.74
A-MEM0.5803,6362.26
MemoryOS0.5533,27632.68
Nemori0.5901,0442.59
MAGMA0.7001,4041.47
Jev-Mem0.7771580.93

Three headline claims: a 0.077 absolute (11.0% relative) quality gain over MAGMA, the strongest baseline; a 6.6× speedup in memory construction over Nemori, the fastest competing system (158 s vs 1,044 s); and 36.7% lower query latency than the fastest baseline (0.93 s vs 1.47 s).

The category breakdown is more interesting than the overall score. Jev-Mem leads four of five question types — multi-hop (0.623), open-domain (0.618), single-hop (0.802), and adversarial (0.962). MAGMA keeps the temporal crown (0.650 vs 0.637). But the adversarial number deserves a pause: questions designed to tempt an agent into misremembering score 0.962 with Jev-Mem versus 0.205 for full-context. A memory system that is more robust to manipulation than simply stuffing the whole history into the window is exactly what you want if agents are going to hold long-lived, adversarial-adjacent contexts — think coding agents ingesting untrusted repos, or assistants reading your email.

Context and caveats

Jev-Mem lands in a crowded field. A-MEM, MemoryOS, Nemori, MAGMA, Mem0, and a small industry of vector-store “memory” products are all racing to own the agent-memory layer, and 2026’s benchmark literature has grown appropriately skeptical — recent surveys complain of benchmark saturation, metric validity problems, and systems that score well but burn 26,000 tokens per query. Against that backdrop, Jev-Mem’s contribution is not “a better memory” so much as “a better governance model for memory”: decisions that used to cost a frontier-model generation now cost a classification.

The honest caveats: LoCoMo is one benchmark, and LLM-as-a-Judge has known biases. The evaluation pairs the controller with a single answer model (GPT-4o-mini), so backbone sensitivity is untested in the reported results. And the GitHub README is admirably direct that the published run commands “do not reproduce the full paper evaluation by themselves.” The paper’s claims are the authors’ measurements, not yet an independent replication.

Still, the direction is right. Inference economics increasingly favor small, non-generative decision models for the 95% of agent work that is routing and gating — a thesis the Jev ecosystem has been building all month, and one that competitors like Laya are now validating from the open-source side. Jev-Mem is the clearest demonstration yet that “System One” models can take over a whole subsystem of the agent stack, not just a classifier slot.

If you run persistent agents, the 158-second build time and sub-second queries are the numbers to benchmark your own memory layer against. The paper, code, and demo are all live today.

Sources