Synced from Zotero on 2026-07-26 19:05 UTC · key
K7234C2E
Authors: Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, Adrià Garriga-Alonso
Published: 2023-10-28
Type: preprint
Collections: Papers
Links: View on Zotero ↗ · DOI ↗ · Source PDF ↗
Abstract
Through considerable effort and intuition, several recent works have reverseengineered nontrivial behaviors of transformer models. This paper systematizes the mechanistic interpretability process they followed. First, researchers choose a metric and dataset that elicit the desired model behavior. Then, they apply activation patching to find which abstract neural network units are involved in the behavior. By varying the dataset, metric, and units under investigation, researchers can understand the functionality of each component.