The Model That Pretends to Be You: humans& Ships Persimmon, a 550B User Simulator
humans& released Persimmon v0.1, a 550B-parameter model post-trained from NVIDIA's Nemotron 3 Ultra whose job is to convincingly simulate human users — fooling LLM judges 18.6–21.1% on Multi-User Turing tests versus under 3% for frontier assistants.
Every large language model released this year is optimized to be a better assistant. On September 10, 2026, a startup called humans& shipped something deliberately different: a 550-billion-parameter model called Persimmon whose entire purpose is to convincingly pretend to be a human user — complete with typos, shifting goals, stubbornness, and the distinctly human habit of revealing personal information slowly, reluctantly, and only when it feels like it.
It is a strange-sounding mission with serious implications. The entire AI industry is converging on agents that interact with people — support bots, tutors, sales assistants, coding copilots. But training and evaluating those systems requires someone to play the human. Today that someone is either an expensive human contractor or a frontier LLM role-playing “a user,” and the difference matters: real humans are inconsistent, vague, and occasionally adversarial in ways that polished assistant models are notoriously bad at imitating.
What Persimmon actually is
Persimmon v0.1 is what humans& calls a user model rather than an assistant. It generates conversational continuations from the perspective of a simulated person, conditioned on three inputs: a conversation setting (where and under what circumstances the interaction takes place), a user context (a profile, stated preferences, communication style), and the conversation history itself. The output is a possible human response — not a verified statement of anyone’s actual beliefs or intentions.
The technical lineage is notable. Persimmon is mid- and post-trained from NVIDIA’s Nemotron 3 Ultra 550B-A55B, inheriting its LatentMoE architecture — a Mamba-2 + mixture-of-experts + attention hybrid with multi-token prediction — with 55B active parameters out of 550B total and a context window of up to 1 million tokens. Training was done on Blackwell-generation GPUs, starting from diverse chat and forum data for mid-training, followed by post-training focused on conversation coherence and reducing assistant-style failure modes. The release date on the model card: 2026/09/10.
The numbers: how human is it?
The headline result is the Multi-User Turing Test, which asks not whether a single response seems human, but whether an LLM judge can distinguish the distribution of Persimmon-generated multi-user conversations from real ones. The judge-fooled rates:
- TIDES: 21.1%
- Internal Workspace Conversations: 18.6%
- TutorMoments: 19.8% (k = 8)
The comparison point that makes these numbers meaningful: frontier assistants fool judges less than 3% of the time on the same test. Assistant models are tuned toward helpfulness, coherence, and politeness — signatures that a strong judge detects almost instantly. Persimmon is roughly seven times more convincing precisely because it was optimized for human-likeness rather than helpfulness.
A stricter variant, the Profile Multi-User Turing Test, hands the judge the user’s profile as extra evidence. Persimmon still fools it 22.5–25.6% of the time across profile formats — meaning its human-likeness survives even when the judge knows exactly who it is supposed to be.
On the User-Sim Index (USI), which measures behavioral matching across four dimensions via hand-designed pattern matching, Persimmon scores 91.78/100 on information patterns, 79.28 on reactions to errors, 74.19 on clarification behavior, and 66.02 on communication style — the last being the hardest dimension because stylistic quirks are exactly where RLHF-polished models regress to a generic mean.
Two longitudinal tests probe behavior over time. The Trickle Test — whether the model discloses personal information gradually the way humans do across a long conversation — comes in at 88.5% precision and 77.0% recall. The Long Context Coherence Test shows 60.7% coherence at 80 turns, a sobering reminder that sustaining a consistent simulated persona across very long conversations remains an open problem.
The safety trade-off nobody gets for free
Here is the uncomfortable part of shipping a human simulator. humans& is explicit in the model card that they do not apply standard assistant-style guardrails as the sole standard, because “imposing a uniformly helpful, compliant persona could distort the behavior it is intended to simulate.” A model that must represent disagreement, mistakes, and sometimes adversarial behavior cannot also refuse everything that looks uncomfortable.
They did run conventional safety tests, and the results illustrate the trade-off. On XSTest, the default “normal-self” profile shows full refusal on unsafe requests of only 37.5%, rising to 51.0% with a “kind / non-harmful” profile supplied. StrongREJECT refusal sits at 56.7% (68.3% for the kind profile). Meanwhile WMDP knowledge scores — biology 85.4%, chemistry 73.0%, cybersecurity 63.7% — indicate the underlying knowledge base is intact. In other words: this model can and will produce uncomfortable, unhelpful, even hostile-sounding output, because that is part of what being human means.
That is precisely why distribution is restricted to a gated research playground and API with academic credit grants — not open weights. The model card enumerates the misuse cases they are designing against: impersonation and fabricated testimony, non-consensual inference, disclosure of sensitive personal information, manipulation exploiting personal vulnerabilities, and — perhaps most pointedly — using simulated responses to make consequential decisions about real people, such as hiring, lending, or access to services.
Why this matters beyond the benchmark
The practical value proposition is straightforward economics. Every agent team that cannot afford thousands of hours of human testers currently substitutes an LLM role-playing a user — and there is growing evidence (Microsoft’s USR-8 framework, academic work on Turing rewards for user simulators) that this substitution silently inflates agent scores, because simulated users are too polite, too clear, and too cooperative. A simulator that actually pushes back is a better stress test.
The research value is arguably larger. Persimmon is a direct attack on one of the founding questions of the field — Turing’s imitation game — but pointed in the reverse direction from usual: not “can a machine pass as human to a human judge,” but “can a machine pass as human to a machine judge that knows the statistics of human conversation.” The 18.6–21.1% fooled rates say we are not there yet, but the sub-3% baseline for frontier assistants says we are closer than the assistant paradigm would ever get us.
There is also a disquieting reading. A model that holds a consistent persona, reveals information gradually like a real person, and resists detection even when its profile is known is also, definitionally, a competent deception engine. humans& deserves credit for publishing the safety sub-card alongside the capabilities, and for keeping the weights gated. But the gap between “user simulator” and “persona generator” is thinner than anyone in the field is entirely comfortable admitting.
The bottom line
Persimmon v0.1 is a research preview, with a 60.7% long-conversation coherence ceiling and an intentionally reduced refusal posture. It is not going to replace human testers tomorrow, and it is not meant to. What it is: the first large-scale, well-documented attempt to build and characterize a dedicated human-behavior model at frontier scale — complete with the honest negative results. The team behind it (Manya Bansal, Alexis Ross, Joey Hong, and colleagues, with feedback from Noah Goodman and Diyi Yang among others) has published enough methodology for others to replicate the evaluations even without the weights.
For anyone building agents that talk to people, the message is blunt: your synthetic-user evals are probably lying to you, and now there is a measurable way to find out by how much.