You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Accuracy vs. KV budget across MATH-500, AIME24, AIME25, and DFS Memory Retention benchmarks. TriAttention consistently outperforms R-KV and SnapKV across all budget levels. On the Recursive State Query benchmark, TriAttention performs comparably to Full Attention up to depth 16, while R-KV shows catastrophic degradation.
Motivation: Q/K Concentration
Pre-RoPE Q/K vectors concentrate around fixed centers, enabling trigonometric modeling of attention patterns. This structure is stable across positions and input contexts.