huggingface.co web signal

EngramEdit Posts 97% Generalization Editing N-Gram Memories

DeepSeek Fine-tuning ai-research llm

TL;DR

  • On CounterFact, EngramEdit reports 99.5% efficacy and 97.0% generalization versus MoEEdit's 67.1% generalization.
  • On ZsRE, EngramEdit reports 97.3% efficacy and 93.7% generalization, with 76.4 utility versus UnKE's 71.9.
  • After 5,000 sequential CounterFact edits, EngramEdit retains over 96% of pre-edit mean F1 across six general-ability tasks.

A new preprint on Hugging Face from Hong Kong Polytechnic University, Hangzhou Diagens Biotechnology, and the University of Science and Technology of China reports 99.5% efficacy and 97.0% generalization on CounterFact, versus 67.1% generalization for the MoEEdit baseline it compares against. The method, EngramEdit, is pitched at conditional-memory LLMs that look up learned embeddings by input n-gram, with DeepSeek Engram cited by name as the architecture family.

The authors frame the problem plainly in the abstract: "different expressions activating different n-gram embeddings." Their response is to compute target memory representations across multiple phrasings of the same fact, jointly update the shared n-gram embeddings, and penalize changes to frequently reused ones. On ZsRE the paper reports 97.3% efficacy and 93.7% generalization, with utility of 76.4 against UnKE's 71.9. On multi-hop MQuAKE under chain-of-thought prompting the paper claims "nearly 3× the accuracy of the strongest baseline," though under standard prompting it says EngramEdit and plain fine-tuning come out similar.

The durability number is the one editing papers usually fail: after 5,000 sequential CounterFact edits, EngramEdit "retains over 96% of pre-edit mean F1 across six general-ability tasks," according to the abstract. One caveat sits inside the paper's own ablation: "Over 92% of new errors on unrelated queries involve activation of updated embeddings," meaning the regularizer limits collateral damage but does not eliminate it. The experiments themselves are run on LongCat-Flash-Lite, not on a DeepSeek checkpoint, which is worth flagging given the headline framing. It arrives in a crowded week for the subfield; our tracker already logs 57 fine-tuning stories in the last 90 days.