When AI Art Has No Author: MIT's 'Attribution Decay' Finding Complicates the Copyright Debate
MIT CSAIL researchers show that as diffusion models scale, individual training images — and even an entire artist's body of work — lose measurable influence on outputs, a phenomenon they call 'attribution decay' with deep implications for copyright, fair use, and artist compensation.
One of the most heated arguments in the AI copyright wars rests on a simple intuition: every AI-generated image is secretly a collage, and if you look hard enough at the training data, you’ll find the picture — or the artist — it came from. A new study from MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL), published August 18 in Nature Communications, suggests that intuition is often simply wrong. And the way the researchers proved it may matter as much as the finding itself.
Zheng Dai, a former MIT CSAIL researcher who led the study, and David Gifford, an MIT professor and CSAIL principal investigator, measured what happens to a diffusion model’s output when specific training data is removed — not approximated, not estimated, but actually deleted and its influence fully eliminated. Their result: the larger the training dataset, the less any individual image matters to what the model generates. At sufficient scale, some generated images cannot be traced to any single training source at all. The authors call this phenomenon attribution decay.
The counterfactual problem
The core question the researchers asked is elegantly counterfactual: what would the model have generated if a particular image had never been in the training set? Answering that directly is brutal. The naive approach — remove one image, retrain the entire model from scratch, regenerate the output, and repeat for thousands or millions of images — is computationally prohibitive. Previous influence-tracing studies got around this by using approximations to estimate how much any one source image contributed to an output.
Dai and Gifford rejected approximations. “All previous methods were approximate,” Gifford said. “They really could not absolutely show that deleting individual things did not change the output. This paper introduces the first method that is absolute. You’re actually deleting the inputs and deleting all influences of the inputs. This is the first exact method for doing large-scale deletion efficiently and showing that the results don’t change.”
Their solution is an architecture they call a diffusion ensemble. Instead of one model trained on the whole dataset, they trained multiple model components on carefully overlapping subsets of the data. Switching off the components that saw a particular image, person, or artist effectively removes that source’s causal influence without retraining anything from scratch. Generate the image again under identical controlled conditions, and you can observe — exactly — whether the removed data mattered.
The logic of attribution, in Dai’s telling, is almost axiomatic: “If you take away a piece of data and the output of the model doesn’t change, then that piece of data didn’t affect the output. So it doesn’t make much sense to attribute the output to that piece of data. And if you then do this one at a time for every other piece of data and find that the output doesn’t change for any of them either, then it doesn’t make much sense to attribute the output to any one of them.”
What they actually tested
The team trained 24 diffusion ensembles on subsets of seven publicly available image datasets, with training sets ranging from just 256 images to more than 162,000. They benchmarked the ensembles against 24 conventional diffusion models trained on the same data, and the images “came out looking about as good by standard measures,” according to press materials — meaning the ensemble trick didn’t degrade output quality.
The pattern that emerged was consistent: the more the training dataset grew, the less any individual example mattered. Crucially, attribution decay wasn’t limited to single pictures. In some experiments the researchers treated every photograph of one person as a single unit; in others, every artwork by a particular artist. The decay held. In one striking illustration, a model trained on public-domain artwork by 744 artists generated nearly identical images regardless of which single artist was removed from the training set.
The researchers are explicit that this is not a get-out-of-jail card for AI companies. Models can and do memorize and reproduce particular training examples, and attributable outputs can occur even at enormous dataset scales. Attribution decay describes a statistical tendency, not a guarantee.
Why it matters
The legal and policy implications are considerable. Gifford frames the results as bearing directly on whether model outputs are merely derivative works. “One way to think about this is that these models are creative,” he said. “They are not simply copying what they are fed, but creating brand new outputs. If those outputs have nothing to do with any individual piece of training data, that raises questions about fair use, about whether the outputs are themselves copyrightable as novel works, and about how authors get compensated when what comes out of a model isn’t attributable to anything on the internet.”
That last clause is the thorny one. Much of the pending litigation and proposed licensing regimes assume a traceable chain from output back to specific training works. If, at scale, that chain frequently dissolves — not because the evidence is hidden, but because no single source carries enough causal weight — then both infringement arguments and compensation mechanisms built on per-work attribution sit on shaky empirical ground. The AlphaGo comparison is instructive: Move 37 depended on everything AlphaGo had learned about Go, yet it wasn’t a retrieval of any human move. Likewise, an AI image can depend on a vast body of human-created work without depending strongly on any single creation.
There is also a constructive flip side. An exact, efficient method for deleting a data point’s influence from a trained model is exactly what opt-out requests, privacy regulations, and court-ordered removals have long needed but lacked. The diffusion ensemble technique is a research prototype, but it demonstrates that “unlearning” at scale can be done exactly rather than approximately — a meaningful data point in its own right.
The honest caveats
The study stops well short of declaring AI creative in the human sense. Humans bring intention, imagination, emotion, and artistic judgment; diffusion models bring pattern synthesis at scale. The finding also doesn’t exonerate training on copyrighted works without permission — whether ingestion itself is fair use is a separate question from whether outputs are traceable, and the study addresses only the latter. And the experiments used datasets up to ~162,000 images, while frontier models train on billions; the trend line points one way, but the largest scales remain extrapolation.
Still, in a debate dominated by strong claims and thin evidence, an exact measurement is a rare thing. Attribution decay gives both sides of the copyright fight something they haven’t had: a rigorous, reproducible handle on what “the source of an AI image” actually means — and sometimes, the answer is that there isn’t one.