Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Valeria Ruscio

Where does Absolute Position come from in decoder-only Transformers?

Jun 04, 2026

Valeria Ruscio, Umberto Nanni, Fabrizio Silvestri

Abstract:RoPE-trained transformers distinguish absolute position in their attention patterns, even though RoPE encodes only relative offsets in the inner product. We trace this leakage to two architectural components, The causal mask is responsible for the first: its per-query softmax denominator depends on the absolute query position by construction. The residual stream supplies the second. Under causal attention the activation at position $0$ attends only to itself and runs as a closed dynamical system from the embedding of the token at that position; downstream attention reads this trajectory through sink-reading heads. Both components appear in all three architectures we study, in architecturally specific balance: NTK scaling suppresses the residual-stream component, sliding-window attention allows it to accumulate with depth, and standard RoPE sits between. Replacing the \texttt{BOS} embedding before the forward pass removes $40\%$ of the residual-stream component at early queries. Attention sinks are token-anchored stabilizers that pass forward a deterministic fingerprint of the token at position $0$, constant across inputs when that token is the auto-prepended \texttt{BOS} and varying with it otherwise.

Via

Access Paper or Ask Questions

The Phenomenology of Hallucinations

Mar 14, 2026

Valeria Ruscio, Keiran Thompson

Abstract:We show that language models hallucinate not because they fail to detect uncertainty, but because of a failure to integrate it into output generation. Across architectures, uncertain inputs are reliably identified, occupying high-dimensional regions with 2-3$\times$ the intrinsic dimensionality of factual inputs. However, this internal signal is weakly coupled to the output layer: uncertainty migrates into low-sensitivity subspaces, becoming geometrically amplified yet functionally silent. Topological analysis shows that uncertainty representations fragment rather than converging to a unified abstention state, while gradient and Fisher probes reveal collapsing sensitivity along the uncertainty direction. Because cross-entropy training provides no attractor for abstention and uniformly rewards confident prediction, associative mechanisms amplify these fractured activations until residual coupling forces a committed output despite internal detection. Causal interventions confirm this account by restoring refusal when uncertainty is directly connected to logits.

Via

Access Paper or Ask Questions

Beyond position: how rotary embeddings shape representations and memory in autoregressive transfomers

Oct 23, 2024

Valeria Ruscio, Fabrizio Silvestri

Figure 1 for Beyond position: how rotary embeddings shape representations and memory in autoregressive transfomers

Figure 2 for Beyond position: how rotary embeddings shape representations and memory in autoregressive transfomers

Figure 3 for Beyond position: how rotary embeddings shape representations and memory in autoregressive transfomers

Figure 4 for Beyond position: how rotary embeddings shape representations and memory in autoregressive transfomers

Abstract:Rotary Positional Embeddings (RoPE) enhance positional encoding in Transformer models, yet their full impact on model dynamics remains underexplored. This paper studies how RoPE introduces position-dependent rotations, causing phase shifts in token embeddings that influence higher-frequency components within the model's internal representations. Through spectral analysis, we demonstrate that RoPE's rotation matrices induce oscillatory behaviors in embeddings, affecting information retention across layers and shaping temporal modeling capabilities. We show that activation functions in feed-forward networks interact with RoPE-modulated embeddings to generate harmonics, leading to constructive or destructive interference based on phase alignment. Our findings reveal that phase alignment amplifies activations and sharpens attention, while misalignment weakens activations and disrupts focus on positional patterns. This study underscores the importance of frequency components as intrinsic elements of model behavior, offering new insights beyond traditional analyses.

Via

Access Paper or Ask Questions

Attention-likelihood relationship in transformers

Mar 15, 2023

Valeria Ruscio, Valentino Maiorca, Fabrizio Silvestri

Figure 1 for Attention-likelihood relationship in transformers

Figure 2 for Attention-likelihood relationship in transformers

Figure 3 for Attention-likelihood relationship in transformers

Figure 4 for Attention-likelihood relationship in transformers

Abstract:We analyze how large language models (LLMs) represent out-of-context words, investigating their reliance on the given context to capture their semantics. Our likelihood-guided text perturbations reveal a correlation between token likelihood and attention values in transformer-based language models. Extensive experiments reveal that unexpected tokens cause the model to attend less to the information coming from themselves to compute their representations, particularly at higher layers. These findings have valuable implications for assessing the robustness of LLMs in real-world scenarios. Fully reproducible codebase at https://github.com/Flegyas/AttentionLikelihood.

Via

Access Paper or Ask Questions