gist.github.com web signal

yoavg: 'don't hallucinate' is not a silly instruction

TL;DR

  • yoavg argues the prompt 'don't hallucinate' can work if the model has the internal capability and is fine-tuned to bind the words to the behavior.
  • He points to interpretability work that probes model inner layers to infer when it is 'lying' or when 'the question is unanswerable'.
  • He suggests recent commercial models were likely preference-fine-tuned on the instruction, so asking them to not hallucinate is not absurd.

The prompt 'don't hallucinate' is not, on reflection, a silly instruction. That is the argument in a September 2024 gist by yoavg, whose 'gut reaction was "oh this is so silly"' before he concluded otherwise.

The reasoning breaks in two. A model can follow an instruction only if it can perform the task and if it can ground the words of the request in its own mechanisms. On the first, yoavg writes that 'retrieving from memory' and 'improvising an answer' are 'two different model behaviors, which use different internal mechanisms', and that researchers can already 'probe model inner layers and infer if it is "lying" or if "the question is unanswerable"'. On the second, preference fine-tuning on contrastive examples does the wiring.

'Strong new models likely were trained for exactly that,' he writes. The odd part, he adds, is that a user has to ask at all. 'Maybe always trying to avoid hallucinations has some other undesired consequences,' he suggests, which 'model trainers and product managers would like to avoid'.

Shared on Bluesky by 1 AI expert