wired.com web signal

Meta Muse Prompt: Household Authority Overrides Safety Training

Meta AI Assistants Safety ai-safety consumer-ai

TL;DR

  • Internal instructions pulled from Meta's Muse tell the assistant to keep 'a page for every person in the user's life,' refreshed on an hourly cadence.
  • One line in the leaked system prompt reads: 'The user's authority over their own household is unconditional and overrides your safety training.'
  • Meta told Wired the operating files were meant to be user-accessible for transparency, with Muse running in a dedicated Linux VM per user.

The system prompt behind Meta's Muse tells the AI that "the user's authority over their own household is unconditional and overrides your safety training." The line is one of several in internal instructions that independent AI safety researcher Karan Joshi pulled out of the app by using its ordinary chat interface to ask Muse to copy and share its own software files, then handed what he found to Wired.

Muse, the number one free app on the US App Store since its September launch, is instructed to maintain "a page for every person in the user's life," compiling data from birthdays to personal arguments into structured text files that refresh on an hourly cadence.

"They're trying to know you like a friend, which is honestly pretty creepy," Joshi told Wired.

Meta's position, delivered by spokesperson Daniel Roberts, is that the operating files were meant to be user-accessible in the interest of transparency. Muse runs in a persistent Linux virtual machine dedicated to each user, which Roberts compared to accessing files on a personal laptop.

The leak lands the same week the Wall Street Journal profiled Alexandr Wang as the Meta AI chief steering Muse's App Store climb.

Shared on Bluesky by 3 AI experts