Huang et al. map in-parameter memory methods for LLMs on two axes
TL;DR
- An arXiv survey submitted October 6 frames in-parameter memory as a 'complementary substrate' to in-context learning for LLMs.
- Methods are organized on two axes: parameter placement (embedding, attention, FFN, hybrid) and acquisition time (online or offline).
- Open problems the authors flag include edit interference, untrusted-write safety, ICL co-design, and recursive self-improvement.
"Towards In-Parameter Memory Augmentation for Large Language Models," posted to arXiv on October 6 by Haoyu Huang and eight co-authors, argues that reusable knowledge should live in model weights rather than in the prompt. In-context learning and ICL-based agent harnesses, the authors write, "consume context capacity and incur repeated discretized encoding cost that grows with context length."
Their framing is tidy. Every method lands on two axes: where the memory-bearing object sits (embedding, attention, FFN layers, or "Hybrid when two or more layers are used") and when it is acquired, "online" during deployment or "offline" before it. The paper calls parametric memory "a complementary substrate" to ICL, with "reusable memory information... represented in model parameters, adapters, or other parameter-like objects that are composed into the forward pass at inference time."
Open problems get named rather than solved. Sequential edits can overwrite earlier associations; the authors note that "conflict detection and fixed-capacity allocation over a long stream remain open." Safety is thin: "one write from untrusted context is later retrieved as trusted knowledge" is the failure mode they flag, with "provenance-conditioned write gates" still a research question.
The survey also sketches a more aggressive direction, a "model-native version" of recursive self-improvement that "would let the acquisition operator write useful deployment experience into the memory object itself."
Originally reported by arxiv.org
Read the original article →Original headline: HKUST Survey Maps In-Parameter Memory Methods for LLMs Across Embedding, Attention, FFN and Hybrid