Transformer 회로에 대한 수학적 프레임워크를 구축한 논의
이 글에서는 어텐션만 사용하는 0~2층 Transformer의 계산 경로를 분해하여, 기계적 해석 가능성을 위한 프레임워크를 설명합니다. 특히, QK 회로를 통해 어텐션 헤드가 정보를 어떻게 결정하는지를 다루며, 이러한 분석이 모델의 동작을 이해하는 데 어떻게 기여하는지를 보여줍니다.
Discussion on a mathematical framework for Transformer circuits.
This article explores a mathematical framework for mechanical interpretability of 0-2 layer Transformers that use only attention. It breaks down the computational path from tokens to output logits and focuses on how the QK circuit determines where attention heads draw information from, highlighting its contribution to understanding model behavior.