RAM-Net Optimizes Linear Attention with Sparse Addressing
New architecture reduces token interference via independent slot addressing, cutting per-step state access by 8x compared to Mamba2.
ImportânciaMaterialEvidênciaE2 não replicadaTratamentoRápido
RAM-Net proposes a sparse address-based access mechanism to address the degradation of long-range fine-grained recall caused by shared states in linear attention.
Traditional linear attention superimposes information from distinct tokens within a fixed-size recurrent state, creating inter-token interference. RAM-Net organizes the state as an array of independent slots and uses an Address Decoder to map each key or query to a sparse address, selecting only a small subset of slots to write to or read from at each step.
Authors report that this design achieves the lowest perplexity with competitive commonsense reasoning and outperforms strong baselines on fine-grained long-range retrieval. A key efficiency metric is accessing fewer state elements per step than all baselines, e.g., 8x fewer than Mamba2.
These are first-party benchmark results from a paper accepted at NeurIPS 2026; independent reproduction has not yet been reported.