Multiple readings
Concern & critique · 1
Questions & unknowns · 1
Ryan Moulton: They are sampled token by token though. I can imagine the "modal token but low probability phrase" bias affecting this.
Pwnallthethings: AIUI, LLMs are not doing linear prediction as a direct markov model. The multiple layers are reducing input sequences of words into (for lack of a non-antropomorphized term) "concepts". The next to…
Open the thread →