Johns Hopkins: 4 chatbots weaken replies to women's phrasing
TL;DR
- GPT-4, Llama, Gemma and Mistral all returned shorter, less formal, lower-grade-level replies when prompts used hedges, tag questions or collective 'we' phrasing.
- Swapping a male or female name into the sign-off had virtually no effect; the models reacted to the writing's register, not the stated identity.
- The paper 'It's How You Ask' will be presented at the Oct. 6-9 Conference on Language Modeling in San Francisco.
Feed the same email request into ChatGPT twice, hedge one version with "maybe" and sign it with a collective "we," and the response comes back shorter, less formal, and pitched at a lower grade level. That is what Johns Hopkins researchers report after running four models (GPT-4, Llama, Gemma and Mistral) against workplace prompts spiked with the hedges, tag questions and collective phrasing more common in women's writing.
"If you prompt a model to write an email you're going to send to someone else at your company, and you're using language features that women more commonly use, you'll get back a response that's less complex, at a lower grade level, and less formal," said senior author Anjalie Field, a Johns Hopkins computer scientist who studies ethics and discrimination in AI.
Swapping the sign-off name did nothing. Male or female, the models responded to how the prompt read, not to whose name was on it. The researchers, whose paper "It's How You Ask: Gender-Associated Linguistic Bias in LLMs" is set for the Oct. 6-9 Conference on Language Modeling in San Francisco, warn the gap could widen as people move to voice systems, since the linguistic patterns are largely unconscious and difficult to change.
"The companies need to fix the models, rather than putting all of the burden on the user," lead author Katherine Van Koevering said.
The Hub write-up gives no percentages, no grade-level delta, and no per-model breakdown.
Shared on Bluesky by 1 AI expert
-
New press release about our recently released COLM paper, work by Katherine Van Koevering! hub.jhu.edu/2026/09/21/a... Check out the full paper here: arxiv.org/abs/2608.13328
View on Bluesky →
Originally reported by hub.jhu.edu
Read the original article →Original headline: AI might be making women sound bad at work