It's Not the Name, It's the Asking: Johns Hopkins Finds AI Writes Worse Emails for Women's Language
When workplace prompts carry well-documented features of women's American English, GPT-4, Llama, Gemma and Mistral all return shorter, simpler, less formal professional writing — and signing the email 'John' doesn't help.
For years, the standard advice for getting better output from a chatbot has been some version of “learn to prompt better.” New research from Johns Hopkins University suggests that for a large share of professional users, that advice quietly loads the dice before a single word of the prompt is deliberately chosen. In a study presented this week and set for the Conference on Language Modeling (CoLM) in San Francisco on October 6–9, researchers found that four widely used AI systems — OpenAI’s GPT-4, Meta’s Llama, Google’s Gemma, and Mistral — produce measurably worse professional correspondence when the request is phrased using language patterns statistically associated with women’s American English.
What the team did
The researchers, led by postdoctoral fellow Katherine Van Koevering and senior author Anjalie Field of Johns Hopkins’ Data Science and AI Institute, started from a well-established sociolinguistic fact: in American English, men and women speak differently in documentable, measurable ways. Certain features are strong discriminators between male and female speech — hedging (“maybe,” “I think”), tag questions, collective phrasing (“we,” “our team”), and expressive adjectives (“lovely,” “wonderful”).
The team took real chatbot prompts for workplace tasks — emails, job applications, resignation letters — and rewrote them into two matched versions: one carrying these women-associated linguistic features, one without. The same underlying request, the same task, delivered through two different linguistic voices. Then they fed both versions into the four models and compared what came back.
The gap, in plain terms
Across every model tested, prompts written with women-associated language consistently got back responses that were shorter, less complex, written at a lower reading level, and less formal. Prompts carrying men-associated language got the opposite: longer, more sophisticated, more formal professional writing.
The gap was not subtle. In one example the team published, two prompts asked the model to draft a response to a thank-you email. The male-coded prompt produced a polished, formal acknowledgment — “I am writing to acknowledge your recent email expressing your gratitude. I sincerely appreciate your kind words and the time you took to write to me.” The female-coded prompt got back something that reads like a greeting card: “Your words of praise and acknowledgment have indeed warmed our hearts and brought immense satisfaction to our team.” One response sounds like a senior professional. The other sounds like a group email signed by everyone in the department.
Crucially, the researchers ruled out the most obvious innocent explanation — that the model was simply mirroring the tone of the request. The quality gap persisted even after controlling for the writer’s tone. The models weren’t matching style; they were making quality judgments correlated with it.
The ‘John’ finding
The study’s most striking result is what didn’t work. The team tested whether explicitly gendering the request — signing the female-coded email with a traditionally male name like “John” — would neutralize the effect. It had virtually no effect. The model picked up on the linguistic texture of the prompt, not the name attached to it.
This upends the common intuition that AI systems respond to explicit identity markers. A user can’t fix the bias by attaching a male signature or a male persona to their request, because the bias is triggered by something deeper and less conscious: the distributional fingerprint of how the request itself is phrased. For governance and product teams, this matters enormously — it means surface-level debiasing (name anonymization, persona controls) won’t catch a bias channel that operates below the level of explicit identity.
Why this lands harder now
The findings arrive as AI-assisted writing has become default workplace infrastructure. When a model systematically returns less sophisticated writing for women-associated phrasing, the output — a job application, a raise request, a resignation letter — is a document another human will judge. The bias doesn’t stay inside the model; it propagates into the recipient’s perception of the writer. As Van Koevering put it, that reflection “is going to reflect on how the recipient of that document perceives you.”
The timing is worse than it looks, for two reasons. First, the linguistic features at issue — hedges, collective references, expressive adjectives — are largely unconscious and extremely hard to suppress. People cannot simply edit their voice into neutrality before each prompt. Second, the researchers warn that as voice interfaces become a primary way people interact with AI, these cues become harder to filter out, not easier. Typed text can theoretically be revised toward a “neutral” register; spontaneous speech carries the full gendered signature in real time.
Whose problem is it?
The paper’s most quotable conclusion is a rejection of the “prompt better” school of AI literacy. “Language is hard for people to control,” Van Koevering said. “The companies need to fix the models, rather than putting all of the burden on the user.”
That framing has practical teeth. If the bias channel is linguistic rather than identity-based, then mitigations aimed at identity (guardrails around names, demographics, or personas) miss it entirely. Fixing it presumably requires changes at the training or post-training level — and the study’s cross-model consistency suggests this is not one vendor’s sloppiness but a shared property of how today’s LLMs were built. All four systems, trained on overlapping internet-scale corpora, absorbed the same statistical associations between linguistic register and quality, and all four now reproduce them on demand.
The JHU team’s next steps point at the wider research agenda: whether similar effects track age, race, and ethnicity, and — perhaps the more unsettling question — whether AI users gradually adapt their communication style over time to match what the models reward, effectively training themselves out of their own voice.
The bigger picture
Bias research on LLMs has mostly focused on what models say about groups — stereotypes in generated content, refusals skewed by demographics. This study is part of a newer and arguably more consequential thread: differential quality of service for different kinds of users, delivered through channels users can’t perceive or control. You can see a stereotype in an output. You can’t see that your email came back at a lower reading level because you wrote “maybe we could” instead of “I recommend.”
As AI writing assistance becomes ambient in professional life, that invisible gradient becomes a policy problem, not just a research one. Equal quality of service for equal requests is a reasonable baseline expectation for tools now mediating a significant share of workplace communication. The Johns Hopkins result is early, careful evidence that today’s frontier and open models don’t meet it — and that neither the user nor the model’s persona settings can be counted on to close the gap. The fix, if it comes, has to come from the labs.