← All posts / Research

Neural Networks Were Secretly Symbolic All Along: Inside the Paper Reconciling AI's Oldest Feud

A new 30-page study from Yale, JHU, NYU, and Microsoft Research shows the vector representations inside MLPs, RNNs, Transformers, and seven open-weight LLMs can be replaced by closed-form symbolic equations with almost no change in behavior.

Neural Networks Were Secretly Symbolic All Along: Inside the Paper Reconciling AI's Oldest Feud

For sixty years, artificial intelligence has been split into two camps. The symbolists — tracing their lineage through Boole, Turing, Chomsky, and Newell and Simon — argued that intelligence means manipulating discrete symbols combined into structured forms: logical formulas, syntax trees, lines of code. The connectionists — from Rumelhart’s 1986 backpropagation revival to today’s frontier labs — bet on continuous vectors, where meaning is a point in a high-dimensional space with no explicit structure at all. Modern AI is the connectionists’ total victory: large language models are, mechanically speaking, giant matrix multipliers over floating-point vectors. And yet they excel at exactly the domains the symbolists claimed as their own — language, arithmetic, logic, and code.

A new paper uploaded to arXiv on August 30, 2026 — “The Emergent Symbolic Structure of Artificial Neural Networks” — proposes a resolution to this paradox, and it has been climbing Hacker News all week. The authors are R. Thomas McCoy (Yale), Paul Soulos (Johns Hopkins), Tal Linzen (NYU), and Paul Smolensky (Microsoft Research) — and the last name matters: Smolensky invented Tensor Product Representations in 1990 precisely as a theoretical answer to how neural nets could encode symbols. Thirty-six years later, his old idea turns out to describe what real networks actually do.

The hypothesis: vectors that secretly hold symbols

The team’s central claim is disarmingly simple: although neural networks are never explicitly designed to represent symbols, they converge on vector representations that implicitly realize symbolic structure anyway. Intelligence, in other words, keeps rediscovering symbols — not because we hard-code them, but because structure is useful and gradient descent finds it.

To test this, the researchers used an analysis framework called DISCOVER (DISsecting COmpositionality in VEctor Representations). The idea is to train a small, fully interpretable “DISCOVER model” to mimic the internal representations of a target neural network — and then to check whether the mimic can be a closed-form equation instantiating a symbolic structure rather than anything resembling a trained network. The formalism chosen is Smolensky’s Tensor Product Representation (TPR), in which a symbolic structure is decomposed into fillers (the elements — say, cats, chase, dogs) each bound to a role (its structural position — subject, verb, object). The role-filler pairs are summed into a single vector, and any element can later be recovered through a linear unbinding operation.

The critical test is brutal: replace the target network’s entire representation-generating process with the TPR equation, then measure whether the network’s downstream behavior survives. If the outputs remain correct, the network’s inner representations weren’t merely correlated with symbolic structure — they were symbolic structure, just wearing vector clothing.

What they found

The results land at every scale the team examined:

1. Emergent symbols across architecture classes. In synthetic sequence-manipulation tasks, DISCOVER found emergent symbolic structure in all three classic neural architectures — multi-layer perceptrons, recurrent networks (GRUs), and Transformers. This is not a Transformer quirk; it appears to be a general property of trained neural systems. In one list-reversal setup, a bidirectional TPR approximation matched the target model perfectly or near-perfectly across all ten training reruns, with the worst single run at 99.98% approximation accuracy. Across architecture-task combinations, the lowest average approximation accuracy was 0.973; every other setting exceeded 0.99.

2. Seven open-weight LLMs, seven hits. The team then analyzed period (”.”) token encodings in seven open-weight models — Gemma-3-27B, GPT-2-XL, GPT-OSS-20B, Pythia-12B, Qwen3-14B, OLMo-2-13B, and Llama-3.1-8B. In all seven, the period vectors were found to compress the structure of the preceding sentence or list, and DISCOVER could closely approximate those encodings with TPRs. Symbolic structure emerges in large-scale systems trained on natural data, not just toy models.

3. Arithmetic, logic, code, and language in one model. Focusing deeply on GPT-OSS, the researchers ran it through four domains long viewed as the symbolists’ home turf: arithmetic, syllogistic logic, code execution, and syntax transformations (passivization, tense reinflection, question formation). Its representations were well approximated by TPRs in all of them — evidence that when an LLM does something symbolic, it is drawing on implicit symbolic structure to do it.

4. Causal interventions, not just correlations. Because a TPR is literally the sum of its role-filler parts, you can edit one part algebraically: subtract the vector for 3rd-position: Z, add 3rd-position: U, and the representation becomes that of the edited list. When the authors performed such edits on real network representations, behavior changed exactly as predicted. Feed a network the clever doctor helped the lawyer, move clever from the “subject adjective” role to the “object adjective” role in its internal representation — and the model then behaves as if it had read the doctor helped the clever lawyer. The identified structure isn’t an epiphenomenon; the model’s behavior depends on it.

5. Systematic generalization. The approximations generalize to novel role-filler combinations: a DISCOVER model trained on sentences where scientist never appeared as the subject still handled scientist-as-subject inputs at test time. The networks compose fillers and roles systematically — a property linguists have long argued is central to human language.

Why this matters

The findings cut against a tempting narrative — that the success of LLMs proves symbolic structure was never necessary. On the contrary: even systems with no symbolic scaffolding built in learn to use symbolic structure on their own, because it is that useful for intelligent behavior. The symbolists lost the architecture war but may have won the conceptual one. As the authors put it, the work offers “a potential way to reconcile longstanding symbolic conceptions of intelligence with the vector-based nature of modern AI.”

For mechanistic interpretability, the implications are concrete. If a model’s representations are implicit TPRs, interpretability has a stronger mathematical toolkit than feature-spotting: one can speak precisely about binding — how a concept is attached to a structural position — which pure “linear representation” accounts (where a sentence is just the sum of its concept vectors) cannot express. That gap is known as the binding problem, and this paper suggests trained networks solve it with role-filler algebra.

There are honest caveats. The results are approximations, not proofs — “largely unchanged” behavior is not identical behavior, and DISCOVER’s role schemes are hypotheses about how a network carves up positions, some of which fit better than others (in the list-reversal experiments, bidirectional roles fit best; left-to-right alone fit poorly). The full code release is still pending employer approval, with a partial codebase on GitHub. And the deep-dive analyses focus on open-weight models — GPT-OSS above all — because white-box access to internal representations is required; frontier closed models remain a black box for this kind of analysis.

Still, the direction is striking. A 1990 theory of how neural nets should encode symbols turns out to be a good description of how 2020s LLMs do. The next time someone declares the symbols-versus-vectors debate settled, the honest answer is now stranger and more interesting: the vectors learned to be symbols.