huggingface.co web signal

Fraunhofer SCAI Paper: 300M-Param PFLM Learns Any Language From Only Synthetic Non-Linguistic Pretraining

Summary

Fraunhofer SCAI's 'Learning to Learn a Language' introduces the Prior-Fitted Language Model (PFLM), a 300M byte-level transformer trained only on samples from a synthetic non-linguistic prior drawn fresh each sequence. With frozen weights it infers the language of a real-text prefix: bits-per-byte fall from 8 to 0.9–2.4 on Wikipedia in six languages at 1M bytes of context, and it compresses six non-text domains below gzip and PPMd. The authors claim it 'has not learned a language' but 'learned to learn one.'