16.79% to 34.97%: The Prefill Experiment That Just Put Qwen 3.8's Training Data on Trial
An independent reasoning-prefill analysis found Qwen 3.8's answer overlap with GPT-5.5 Pro more than doubles when seeded with 1% of GPT's chain-of-thought — the sharpest technical signal yet in the distillation debate.
Two days after the NSA, CISA, and FBI formally accused six Chinese AI companies of running industrial-scale distillation campaigns against US frontier models, an independent researcher has published the sharpest technical signal in that debate so far — and it points at a model nobody named in the advisory.
On September 9, Meng Zhang (GitHub: wsxiaoys), the creator of the Tabby coding assistant, published a quiet follow-up experiment titled “Reasoning prefills on a few open models, v1.1.” By September 10 it sat near the top of Hacker News with over 230 points and 91 comments. The finding inside is simple to state and hard to explain away: when you insert the first 1% of GPT-5.5 Pro’s chain-of-thought into Qwen 3.8 A95B’s reasoning channel and let it finish on its own, its final answers converge toward GPT-5.5 Pro’s answers at more than double the baseline rate.
The experiment
The method is disarmingly cheap. For each of 45 problems — 15 STEM, 15 non-STEM, and 15 synthetic puzzles written by the author and guaranteed absent from any training set — the researcher generated two responses from each target model: an ordinary unprefilled response, and a response whose reasoning channel begins with the first 1% of GPT-5.5 Pro’s own reasoning. The visible answer was always left freely generated. Overlap was then scored as the mean of unigram, bigram, and trigram source recall over the first 100 tokens of the target’s answer.
The headline numbers for the full problem set:
| Model | Unprefilled | GPT-5.5 Pro prefill | Delta |
|---|---|---|---|
| DeepSeek V4 Flash | 27.30% | 26.13% | −1.17 pp |
| Inkling | 19.99% | 20.45% | +0.46 pp |
| Kimi K3 | 31.11% | 35.65% | +4.54 pp |
| Qwen 3.8 A95B | 16.79% | 34.97% | +18.18 pp |
Three of the four open models barely move. Qwen 3.8 more than doubles its convergence. And the breakdown by category makes it more interesting, not less: on STEM problems the jump is +26.99 percentage points (19.26% → 46.24%), on the author’s private synthetic puzzles it is +14.75 points (10.49% → 25.23%). The synthetic-puzzle result matters because those problems cannot have appeared in any scraped corpus — whatever makes Qwen follow GPT’s lead there was learned as a behavior, not memorized as content.
The intuition behind the method: reasoning models tend to complete a begun thought in their own trained voice. If a model was post-trained on another model’s reasoning traces, seeding it with a fragment of that teacher’s thinking should feel like “coming home” — pulling its continuation and final answer measurably toward the teacher’s. A model with an unrelated training history should shrug the prefill off, which is exactly what DeepSeek V4 Flash (−1.17 pp) and Inkling (+0.46 pp) do.
Where the tool came from
The prefill methodology descends from research with a much sharper edge. In August 2026, a team including Maksym Andriushchenko and Alexander Panfilov published “Stealing Reasoning Traces from Proprietary LLM APIs” (arXiv:2608.09867), which documented an architectural vulnerability in how every major provider ships hidden chain-of-thought: encrypted reasoning blocks returned to the client are interchangeable across sessions, users, and models within a provider’s ecosystem. By injecting an encrypted trace from a strong model into a weaker, less-safeguarded model from the same provider, an attacker can force the weaker model to decode the trace verbatim in plaintext — no jailbreak of the capable model required.
The paper demonstrated four attack vectors: circumventing anti-distillation defenses across Anthropic, OpenAI, and Google; mass extraction of private data (367 PII artifacts and 182 credentials recovered from 315,320 reasoning blocks scraped from public repositories); disclosure of hazardous content hidden inside reasoning that the visible output safely refused; and invisible prompt injections embedded entirely within encrypted blocks.
The v1.0 experiment from the same author in early September used Opus 4.8 reasoning as the teacher seed and found the strongest convergence on Kimi K3 — consistent with the paper’s anecdote that “prefilling Kimi-K3’s reasoning with just a few tokens of Opus reasoning measurably shifts its response toward Opus’s.” Qwen barely moved toward Opus in that round. The v1.1 rerun with GPT-5.5 Pro as teacher flipped the picture: Qwen 3.8 is the outlier, and the author’s stated conclusion is careful — “the data suggest that Qwen may have learned from GPT-5.5 Pro, or from a closely related GPT model, rather than from Opus.”
Why the timing turns up the heat
The geopolitical context is what elevates a 45-problem gist into a story worth following. On September 8, the NSA, CISA, and FBI published joint advisory AA26-251A naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI as conducting distillation campaigns that allegedly “form the core — not merely a supplement” of those companies’ AI development strategy. Alibaba — Qwen’s owner — is on that list. The advisory even described the reconstruction of hidden chain-of-thought as an explicit adversary technique.
Independent technical evidence landing 48 hours later, pointing in the same direction with a clean experimental design, is exactly the kind of corroboration that outlives a news cycle. It is also, for now, evidence rather than proof, and the gaps deserve honest statement: the effect is correlational, the sample is 45 problems, the identity of the actual teacher (GPT-5.5 Pro proper, a sibling GPT variant, or some mixture) is unresolved, and high prefill convergence is at least conceivable through convergent post-training recipes rather than direct trace ingestion. The author flags these limits; the HN comment thread stress-tests them further, including the alternative hypothesis that shared benchmark-style training data could produce similar overlap. The counter-consideration is that the synthetic puzzles — unseen by construction — show the effect too.
What it means
For Alibaba, the timing is awkward at minimum. Qwen 3.8-Max, released August 2, scaled to 2.4 trillion parameters and was billed as a new bar for coding and cowork, a genuinely strong open-weight family by every public benchmark. A distillation question mark does not erase the engineering, but it hands regulators, customers, and rivals a concrete, reproducible experiment to cite — and the NSC-adjacent agencies have already shown they will act on this category of allegation.
For the research community, the more durable contribution may be methodological. Prefill-convergence is cheap, model-agnostic, and hard to game after the fact: you cannot un-learn a teacher. As one HN commenter put it, “reasoning works as long as there is a consistent latent space representation” — which means the fingerprint a teacher leaves in a student is structural, not cosmetic. Expect this template to be run against every frontier-model pairing that matters, with control models and larger problem sets. The author’s own “what to watch” list starts there: independent replication, ideally with a control model carrying no distillation allegation.
For everyone else, the through-line of September’s AI news is that the industry’s biggest arguments are no longer settled by press releases. A joint federal advisory accuses; an encrypted-blocks exploit exposes; a 45-problem prefill experiment measures. The documents and the data are doing the talking — and this time, one of them is talking about Qwen.
The experiment, including tables and methodology notes, is public on GitHub Gist; the underlying trace-extraction technique is documented in arXiv:2608.09867.