arxiv.org web signal

'Trust No Bot' study finds PII in 48% of GPT translations

TL;DR

  • Personal data appeared in 48% of translation prompts and 16% of code-editing prompts that real users sent to commercial GPT models.
  • The authors argue PII detection tools miss sensitive disclosures like detailed sexual preferences or specific drug use habits.
  • The paper prescribes 'nudging mechanisms' that help users moderate what they share with chatbots at prompt time.

Personally identifiable information showed up in 48% of translation tasks and 16% of code-editing tasks that real users sent to commercial GPT models, according to a preprint on arXiv. The paper builds a taxonomy of tasks and sensitive topics from what it calls naturally occurring conversations between people and chatbots.

The authors, Niloofar Mireshghallah, Maria Antoniak, Yash More, Yejin Choi and Golnoosh Farnadi, describe the leakage as often unexpected. "Personally identifiable information (PII) appears in unexpected contexts such as in translation or code editing (48% and 16% of the time, respectively)," they write.

They also argue that standard redaction tooling misses the more slippery half of the problem. "PII detection alone is insufficient to capture the sensitive topics that are common in human-chatbot interactions, such as detailed sexual preferences or specific drug use habits."

Their prescription is a UX one, not a model one. The paper calls for "appropriate nudging mechanisms to help users moderate their interactions."

The abstract does not name the underlying conversation dataset, the sample size, or which commercial GPT model versions were included.

Shared on Bluesky by 1 AI expert