Paper
On the Role of Attention Heads in Large Language Model Safety
ICLR 2025
Methods
Previous research has shown that there are safety parameters in LLM. This paper focuses on the attention heads.
Ships (Safety Head Important Score)
Ships helps to qualify the impact of a specific attention head on safety. As shown in the formula, it calculates the KL divergence between two output distributions: one with all attention heads and another with a head ablated.
$q_{\mathcal{H}}$ is harmful queries. $\theta_{\mathcal{O}}$ is the parameters of original attention heads. $\theta_{h_{i}^{l}}$ is the parameters of the $i$-th attention head in layer $l$. $\setminus$ is the ablation operation.
How to ablate an attention head? The paper proposes two methods.