Episode 2238
· 20:59
🤗 Upvotes: 37 | cs.CL, cs.AI, cs.LG
Authors:
Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
Title:
Language Models Can Control Their Own Attention
Arxiv:
http://arxiv.org/abs/2609.02737v1
Abstract:
Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to find the few tokens that matter. If the user asks about a previous detail in a 1M-token conversation, global attention layers must scan the full context to generate each token of the reply. A prominent approach mitigates this cost by pre-selecting relevant tokens via lightweight proxy scores, but this extrinsic scoring still incurs O(N) per step. We take an intrinsic approach motivated by the simple question: wouldn't the model already know which parts of the context are relevant? To this end, we introduce Declarative Attention (DA), a protocol that elicits the model to declare where it needs to attend within its chain-of-thought, partitioning generation into three modes:
Listen to Daily Paper Cast using one of many popular podcasting apps or directories.