← All posts / Research

Two Thoughts, One Pass: Researchers Show Transformers Can Superpose Text Streams

A nine-author arXiv paper demonstrates that averaging the embeddings of two documents makes an LLM output a blend of both next-token distributions — an intrinsic architectural property that pretraining erodes and lightweight fine-tuning can restore.

Two Thoughts, One Pass: Researchers Show Transformers Can Superpose Text Streams

A paper posted to arXiv on September 24 carries a title engineered to stop a scroll: “Your Transformer Can Hold Two Thoughts at Once.” Behind the hook is a genuinely strange empirical claim with real implications for how we think about LLM inference — and possibly for how much we pay for it.

The work, led by Pavel Tikhonov, Anton Korznikov and Matvey Mikhalchuk together with Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets and Elena Tutubalina, proposes what they call the Superposition Linearity Hypothesis. The claim: when you linearly combine the input embeddings of two distinct text streams — say, element-wise averaging the token embeddings of document A and document B — a pretrained Transformer does not collapse into noise. Instead, it outputs something close to a superposition of the two individual next-token distributions. Both “thoughts” survive inside the same forward pass.

Why this is surprising

Everything about a modern LLM screams non-linear. Self-attention, MLP blocks, softmax over the vocabulary — none of it looks like the kind of system where you can average two inputs and expect anything sensible to come out. The prevailing paradigm treats inference as a single coherent semantic stream: if you want to process two independent documents, you run two forward passes, or you reach for specialized architectures designed to avoid destructive interference.

Yet prior work (Razzhigaev et al., 2024 — an overlapping group) showed that decoder-only Transformers exhibit strong linear structure in the residual stream: transitions between consecutive layers can often be well approximated by affine maps. The new paper asks the natural follow-up: does that linearity extend all the way to end-to-end input–output behavior? Their answer is yes, to a striking degree. When embeddings from two distinct documents are averaged token-wise and fed to standard pretrained LLMs, the ground-truth next tokens for both streams consistently retain substantial probability mass — frequently landing in the top-10 ranks of the mixed output distribution.

The finding that cuts against the grain

The most philosophically loaded result is buried in Section 2.4. The authors track superposition fidelity across the pretraining trajectory and find it is maximized at initialization — before the model has learned anything — and gradually diminishes as the model optimizes the language modeling objective.

In other words, superposition is not a capability the model acquires. It is an intrinsic property of the Transformer architecture that training actively erodes. The standard pretraining objective, which only ever sees single coherent streams, has no reason to preserve the ability to process mixtures — and so it doesn’t. The authors further observe a strong correlation between geometric linearity in hidden states (measured via feature additivity) and rank preservation under superposition, tying the input–output effect to measurable structure inside the network.

An attention-patching analysis adds nuance: positions whose behavior is predictable survive mixing when the overall attention shape is preserved, while content-carrying positions are far more fragile — and it is exactly those content positions that lightweight fine-tuning brings back.

Restoring linearity on a budget

If pretraining degrades the property, can it be recovered? Yes, and cheaply. The authors apply a self-distillation setup: a student model, initialized from pretrained weights, processes the averaged embeddings, while a frozen copy of the same model acts as teacher, supplying the target distribution — the arithmetic mean of its independent predictions on each stream. The loss is simply the KL divergence between that mixture and the student’s output on the mixed input.

The numbers on Pythia-2.8B are dramatic. Mean KL divergence between the mixed prediction and the analytical target drops from 1.86 to 0.27, and the Superposition Approximation Ratio falls from 0.42 to 0.06 — roughly a sevenfold improvement in how faithfully the mixed forward pass tracks the average of the two independent distributions. Crucially, this costs under 0.025% of the original pretraining dataset size (about 200k steps on a FineWeb subset), applied across Pythia, Qwen and Llama model families. This is not a retrain; it is a nudge.

The hard part: getting the thoughts back out

Fitting the mixed distribution is one thing. Decoding two clean, separate continuations from it is where the paper hits its most interesting wall — the geometric-mean obstruction.

When logits are approximately averaged, the resulting mixed probabilities scale with the geometric mean of the independent distributions. That creates a structural penalty: any token highly probable in stream A but unlikely in stream B gets crushed in the mixture, even when both ground-truth tokens sit reliably in the top-5 after fine-tuning. Naively sampling from the mixed distribution produces semantically incoherent sequences that alternate between the two documents’ tokens.

As a proof of concept, the authors introduce Joint Contrastive decoding, using a small auxiliary model to help disentangle the mixed hidden state back into its constituent streams — enough to demonstrate simultaneous generation of two coherent continuations from a single forward pass, with additional decoding variants (a parameter-free two-head approach, inference-time logit arithmetic) deferred to the appendix. They are candid that fully overcoming the obstruction “remains an open problem for future research.”

Why throughput people should care

Strip away the interpretability angle and there is a bluntly practical pitch. If two streams can be compressed into one forward pass and then disentangled:

  • ~2× inference throughput — two continuations generated for the price of one pass.
  • Halved KV-cache footprint per active stream — multiple streams share a single sequence of mixed vectors, a meaningful lever on the memory bottleneck that dominates LLM serving costs.

The lineage here is real: DataMUX (2022) multiplexed inputs into single representations; Superposed Decoding (2024) mixed draft token embeddings for parallel generation; superposition prompting (2024) accelerated RAG by processing multiple document paths in one pass. What distinguishes this paper is that it doesn’t add mux/demux machinery — it isolates an intrinsic input–output superposition effect already present in standard pretrained LLMs, tracks its lifecycle across training, and shows a minimal recipe for reviving it.

The boundaries

The limitations section is refreshingly specific. Evaluations used short contexts (L ≤ 128 for the analytical experiments, extended to 512), predominantly monolingual corpora, and text-only models. Whether the effect survives at production context lengths, in multilingual settings, or in multimodal architectures — where mixing text tokens with image patches poses geometrically distinct challenges — is entirely unexplored. And of course, a proof-of-concept decoder on TinyStories-scale coherence is a long way from a production serving stack.

Still, as a statement about what Transformers are rather than what we train them to do, the paper lands a clean point: the architecture is more linear — and more multiplexable — than the way we use it suggests. The two-thoughts-at-once trick is sitting there in the weights. Pretraining talks it out of them; a sliver of fine-tuning brings it back.