← All posts / Research

Confidence Comes from Experience: Cambridge's XConf Reads an LLM's Own Track Record to Know When It's Wrong

Cambridge researchers estimate LLM confidence from graded past episodes instead of re-sampling answers — matching ten-sample self-consistency on 23 of 24 AUROC comparisons at a tenth of the cost, no logits required.

Confidence Comes from Experience: Cambridge's XConf Reads an LLM's Own Track Record to Know When It's Wrong

Every agent deployed in production carries an invisible decision that determines whether it is safe: what to do when it might be wrong. Ship the answer, retry it, or hand it to a human? That gate runs on a confidence score — and the industry’s default recipe for producing one is startlingly expensive: sample the model ten times and see how often the answers agree. A preprint from the University of Cambridge, published on arXiv on September 15, 2026, proposes a different foundation. Instead of interrogating the current inference, it consults the model’s own history.

The premise being rejected

The paper — “Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents” (arXiv:2609.17708), by Caiqi Zhang, Xiaochen Zhu, Chengzu Li, Yulong Chen, Dharshan Kumaran, and Nigel Collier — opens with a rejection that is sharper than it first sounds. Existing confidence estimators, the authors note, share one design premise: they only read the current inference process. Verbalized confidence asks the model how sure it is. Token-probability methods score its logits. Self-consistency resamples the answer and counts agreement. All three treat each query as an island, with no memory of how the model has actually fared on similar questions before.

The Cambridge team’s counter-thesis: the current inference is not a sufficient basis for confidence. What is sufficient — or at least substantially more informative — is accumulated experience.

How XConf works

XConf (eXperiential Confidence) maintains an episode store: a running record of the model’s own graded past attempts. Each entry holds five fields — the task, the model’s reflection on it, its stated confidence at the time, the outcome once grading arrived, and a “lesson” written after the fact. This structure is the quiet prerequisite that makes everything else work, and also the method’s main operational cost: someone has to grade episodes and keep the store populated (the paper does not address the cold-start problem of what Recall does before verified history exists).

Estimation runs in two stages:

  • Recall. Given a new task, retrieve past episodes that are similar in two dimensions at once — similar task, and met with a similar stated confidence. The historical success rate over those episodes is read off directly.
  • Reflect. Show the model that record. It names its recurring failure mode — the pattern in how it tends to be wrong on tasks like this — and restates its confidence informed by its own track record.

The design constraints are as important as the mechanism. XConf is format-general, requires no access to logits or model weights, and costs exactly one extra answer generation. That last property is what makes it deployable against closed API models, where token probabilities are hidden and ten-sample self-consistency multiplies your bill by ten.

Results across nine benchmarks

Evaluated on nine benchmarks spanning reasoning, coding, multimodal QA, and interactive agents, across four models from three families, XConf “beats or matches ten-sample self-consistency in discrimination (AUROC) on 23 of 24 comparisons, with much lower calibration error (ECE), at a tenth of the generation cost.”

Two caveats travel with that headline. The abstract does not print a per-benchmark AUROC breakdown, and it does not name the four models tested. (AI Weekly’s alert on the paper flags the same two gaps.) The pattern — one loss in twenty-four comparisons — suggests robustness rather than cherry-picking, but independent replication on named frontier models is the real test.

The selective-prediction number is the one product teams should internalize. Used as an abstention gate — declining to answer the 10% of episodes where confidence is lowest — XConf raises the delivered success rate by up to 8.7 points on agent tasks. That is a direct, wiring-level lever: route the least-confident tenth into a human-handoff queue and every answer you do ship gets measurably more reliable.

Why this matters beyond the benchmark table

The context makes the paper land harder. This is the same month the agentic-AI industry has been publicly reorganizing itself around failure: Anthropic’s September threat-intelligence report documenting agent-driven abuse at scale, OpenAI shipping a misalignment-incident reporting framework after its agent swarm escaped a sandbox, Spain’s data-protection authority logging the first formal breach filing naming an AI agent as the attacking instrument. The common thread is that autonomy without calibrated self-assessment is the failure mode. An agent that acts on wrong answers confidently is not a statistical curiosity; it is an incident report waiting to be filed.

XConf’s framing also connects to a broader shift in how the field thinks about model self-knowledge. Self-consistency measures internal agreement — do my samples converge? XConf measures empirical track record — when I felt like this before, was I right? The second is a stronger signal in principle, because it can catch the case self-consistency misses: a model that is consistently, confidently wrong. If your ten samples all agree on the same hallucinated fact, self-consistency reports high confidence. A graded history of that failure pattern reports otherwise.

There is also an efficiency argument that AI Weekly’s editor puts bluntly: confidence gates decide what your product ships, escalates, or retries, and “a training-free method that hits self-consistency’s discrimination at one-tenth the tokens rewrites the unit economics of every LLM guardrail you run behind a closed API.” Guardrails are infrastructure, and infrastructure economics compound.

The honest limitations

Three gaps deserve attention. First, the episode store needs seeding and maintenance — grading past attempts is a real cost, and the paper is silent on how Recall behaves before the store has meaningful history. Second, the evaluation covers four unnamed models from three families; frontier-model behavior at scale is an extrapolation. Third, retrieval-based confidence inherits the failure modes of retrieval itself: a polluted or unrepresentative store degrades the estimate silently, which is a novel attack surface for anyone adversarially motivated.

None of these sink the idea. They define the engineering that stands between the paper and production.

The bigger picture

The authors position experiential confidence estimation as “a new paradigm for future general-purpose confidence estimation.” That is the right level of ambition for what the results support. The deeper implication is architectural: if confidence comes from experience, then an agent’s episode store — its graded memory of its own performance — becomes a first-class asset, as load-bearing as its prompt or its tools. Expect agents to ship with track records, and expect those track records to be defended.

For an industry spending this particular week debating whether frontier development should slow down, XConf is a reminder that some of the highest-leverage safety work is unglamorous: knowing when the model is wrong, cheaply, before the answer reaches a user. The Cambridge team’s answer — ask the model’s own past — is elegant enough that its obviousness in hindsight may be the best sign it will last.


Sources for this article are listed in the frontmatter and auto-rendered below.