Fine-tuning a model to just remember a stream of facts sounds simple until you watch retention collapse. The long-horizon memorization setting in this paper — 100 query-answer tasks learned by sequential fine-tuning, no replay, no task IDs at inference — drives naive supervised fine-tuning down to 1.2% final retention. The headline isn’t a new mechanism; it’s that no single mechanism survives the horizon, and composing the right ones does.

The HF paper page frames this as memorization, but the real target is any model you keep updating in place instead of re-indexing. If you’re choosing between fine-tuning knowledge in and retrieving it at query time, does 34.9% retention change the math for your slice of facts?

tags: [ research ] [ llm-ops ] [ rag ]