Bldev's Blog

순환도 합성곱도 없다: 트랜스포머가 어텐션만으로 세상을 바꾼 방법

2026. 8. 27.

구조 및 방법론 비교
과제: 기계 번역 (WMT 2014 EN-DE, EN-FR) 및 구성 성분 파싱
기존 접근법복잡한 순환(RNN, LSTM) 또는 합성곱(CNN) 신경망의 인코더-디코더 구조로, 내재적 순차 계산으로 인해 시퀀스 내 병렬화가 불가능하고 장거리 의존성 학습 경로가 긺
본 논문 제안순수 self-attention 메커니즘만으로 구성된 완전 병렬화 가능한 시퀀스 변환 모델로, 임의 위치 간을 상수 횟수 연산으로 연결하여 훈련 속도와 성능 모두 획기적 개선

순환과 합성곱을 완전히 제거하고 attention만으로도 우수한 성능 달성 가능하며, 병렬화·훈련 속도·SOTA 성능 모든 측면에서 획기적 혁신을 입증

비교 항목기존 기준선Transformer (본 논문)근거
순차 연산 복잡도 (최대 경로 길이)[RNN (순환 신경망)] O(n) — 시퀀스 길이에 선형 비례O(1) — 상수 횟수 직결[원문]
임의 위치 간 신호 연결 연산 수 증가율[ConvS2S (합성곱 시퀀스-투-시퀀스)] 선형적(linear) 증가 / ByteNet은 대수적(logarithmic) 증가상수(constant) 연산으로 감소[원문]
WMT 2014 English-to-German 번역 성능 (BLEU)[이전 SOTA (ensembles 포함)] ≤26.4 BLEU (Transformer big 대비 2.0 BLEU 이상 낮음)28.4 BLEU (Transformer big) — +2.0 이상 향상[원문]
WMT 2014 English-to-French 번역 성능 (BLEU)[이전 단일 모델 SOTA] <41.0 BLEU41.0 BLEU (Transformer big) — 새로운 단일 모델 SOTA 수립[원문]
EN-DE 번역 훈련 비용 (시간/자원)[이전 SOTA 모델] 4배 이상의 훈련 비용 소요3.5일 (8 P100 GPU) — 이전 최고의 1/4 미만[원문]
어텐션 헤드 수에 따른 EN-DE 번역 성능 (BLEU)[단일 헤드 어텐션 (h=1)] -0.9 BLEU (최적 설정 대비 하락)h=8 멀티-헤드 (최적 설정)[원문]

1. 연구 배경 및 핵심 문제 의식

2017년 이전까지 자연어 처리의 시퀀스 변환 과제는 순환 신경망, 특히 장단기 메모리와 게이트 순환 신경망이 장악하고 있었다. 인코더-디코더 구조 위에 어텐션 메커니즘을 얹은 이 모델들은 당시 기계 번역의 사실상 표준이었다. 그러나 이 패러다임에는 태생적인 한계가 존재했다.

"Recurrent models typically factor computation along the symbol positions of the input and output sequences. Aligning the positions to steps in computation time, they generate a sequence of hidden states ht, as a function of the previous hidden state ht−1 and the input for position t. This inherently sequential nature precludes parallelization within training examples, which becomes critical at longer sequence lengths, as memory constraints limit batching across examples. Recent work has achieved significant improvements in computational efficiency through factorization tricks [21] and conditional computation [32], while also improving model performance in case of the latter. The fundamental constraint of sequential computation, however, remains."
— 원문 링크

순환 모델의 가장 근본적인 문제는 계산이 본질적으로 순차적이라는 점이다. 각 타임스텝의 은닉 상태는 반드시 이전 상태에 의존하기 때문에, 학습 예제 내부에서의 병렬화가 불가능하다. 시퀀스가 길어질수록 메모리 제약으로 배치 처리마저 제한되어 학습 비용이 급격히 증가한다. 합성곱 신경망 기반의 대안 모델들, 예컨대 바이트넷이나 합성곱 시퀀스-투-시퀀스 모델도 제안되었지만, 두 임의 위치 사이의 신호를 연결하는 데 필요한 연산 수가 거리에 따라 선형 또는 로그적으로 늘어나 장거리 의존성 학습의 어려움을 근본적으로 해소하지 못했다.

"The goal of reducing sequential computation also forms the foundation of the Extended Neural GPU [16], ByteNet [18] and ConvS2S [9], all of which use convolutional neural networks as basic building block, computing hidden representations in parallel for all input and output positions. In these models, the number of operations required to relate signals from two arbitrary input or output positions grows in the distance between positions, linearly for ConvS2S and logarithmically for ByteNet. This makes it more difficult to learn dependencies between distant positions [12]. In the Transformer this is reduced to a constant number of operations, albeit at the cost of reduced effective resolution due to averaging attention-weighted positions, an effect we counteract with Multi-Head Attention as described in section 3.2."
— 원문 링크

Vaswani 등은 이 두 가지 제약, 즉 순차적 계산 구조와 위치 간 거리에 비례하는 연결 비용을 동시에 해결하기 위해 근본적으로 다른 방향을 택했다. 순환과 합성곱을 완전히 제거하고, 오직 어텐션 메커니즘만으로 구성된 새로운 시퀀스 변환 아키텍처, 트랜스포머를 제안한 것이다. 이는 당시의 관행을 전면적으로 뒤집는 시도였으며, 논문의 제목 'Attention Is All You Need'는 이 대담한 선언을 그대로 담고 있다.

"To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution. In the following sections, we will describe the Transformer, motivate self-attention and discuss its advantages over models such as [17, 18] and [9]."
— 원문 링크

2. 핵심 발견 및 실증 분석

트랜스포머의 혁신은 단순한 아이디어 제안에 그치지 않는다. 아키텍처 설계의 세부 선택들, 그리고 이를 뒷받침하는 치밀한 실험들이 함께 논문의 설득력을 만들어낸다. 이하에서는 주요 기여를 세부 주제별로 살펴본다.

2-1. 셀프 어텐션이 순환 연산을 대체할 수 있는 이유

트랜스포머의 핵심 주장은 셀프 어텐션이 순환 레이어보다 장거리 의존성을 더 효율적으로 학습할 수 있다는 것이다. 셀프 어텐션 레이어는 시퀀스 내 임의의 두 위치를 상수 횟수의 순차 연산으로 연결하는 반면, 순환 레이어는 두 위치 사이의 거리에 비례하는 O(n)O(n)의 순차 연산을 필요로 한다. 연결 경로가 짧을수록 역전파 과정에서 기울기 소실 없이 의존성을 포착하기 쉽다.

"As noted in Table 1, a self-attention layer connects all positions with a constant number of sequentially executed operations, whereas a recurrent layer requires O(n) sequential operations. In terms of computational complexity, self-attention layers are faster than recurrent layers when the sequence length n is smaller than the representation dimensionality d, which is most often the case with sentence representations used by state-of-the-art models in machine translations, such as word-piece [38] and byte-pair [31] representations. To improve computational performance for tasks involving very long sequences, self-attention could be restricted to considering only a neighborhood of size r in the input sequence centered around the respective output position. This would increase the maximum path length to O(n/r). We plan to investigate this approach further in future work."
— 원문 링크

계산 복잡도 측면에서도 기계 번역처럼 시퀀스 길이 nn이 표현 차원 dd보다 작은 일반적인 상황에서는 셀프 어텐션이 순환 레이어보다 더 낮은 계산 비용을 갖는다. 다만 매우 긴 시퀀스를 처리할 경우 이 이점이 반전될 수 있으며, 저자들은 이를 위해 출력 위치 중심으로 반경 rr의 이웃만 참조하는 지역화 전략을 추후 연구 과제로 남겨두었다.

"As noted in Table 1, a self-attention layer connects all positions with a constant number of sequentially executed operations, whereas a recurrent layer requires O(n) sequential operations. In terms of computational complexity, self-attention layers are faster than recurrent layers when the sequence length n is smaller than the representation dimensionality d, which is most often the case with sentence representations used by state-of-the-art models in machine translations, such as word-piece [38] and byte-pair [31] representations. To improve computational performance for tasks involving very long sequences, self-attention could be restricted to considering only a neighborhood of size r in the input sequence centered around the respective output position. This would increase the maximum path length to O(n/r). We plan to investigate this approach further in future work."
— 원문 링크

2-2. 멀티 헤드 어텐션과 스케일드 닷 프로덕트 어텐션의 설계

어텐션 함수는 쿼리와 키-값 쌍의 집합을 출력으로 매핑하는 연산이다. 기존의 덧셈 어텐션 방식과 닷 프로덕트 어텐션 방식을 비교했을 때, 작은 차원에서는 두 방식의 성능이 비슷하지만 키 차원 dkd_k가 커질수록 닷 프로덕트의 내적 값이 크게 증가해 소프트맥스의 기울기가 극히 작은 영역으로 밀려나는 문제가 발생한다. 이를 해결하기 위해 스케일드 닷 프로덕트 어텐션은 내적 결과를 1/dk1/\sqrt{d_k}로 나눠 기울기 소실을 방지한다.

"While for small values of dk the two mechanisms perform similarly, additive attention outperforms dot product attention without scaling for larger values of dk [3]. We suspect that for large values of dk, the dot products grow large in magnitude, pushing the softmax function into regions where it has extremely small gradients . To counteract this effect, we scale the dot products by 1dk."
— 원문 링크

단일 어텐션 헤드로 모든 위치를 한꺼번에 바라보면 서로 다른 표현 부분공간의 정보가 평균화되어 세밀한 패턴을 포착하기 어렵다. 멀티 헤드 어텐션은 이 문제를 해결하기 위해 쿼리, 키, 값을 hh개의 다른 선형 투영으로 변환한 뒤 각각 어텐션을 수행하고 그 결과를 이어붙이는 방식을 취한다. 논문에서는 h=8h=8개의 헤드를 사용하며, 각 헤드의 차원을 dmodel/h=64d_{\text{model}}/h=64로 줄여 전체 계산 비용이 풀 차원 단일 헤드 어텐션과 유사하게 유지되도록 설계했다.

"Multi-head attention allows the model to jointly attend to information from different representation subspaces at different positions. With a single attention head, averaging inhibits this."
— 원문 링크

2-3. 위치 인코딩: 순서 정보를 어텐션에 주입하는 방법

어텐션 메커니즘은 집합 연산이기 때문에 토큰의 순서 정보를 자체적으로 담지 못한다. 트랜스포머는 이를 보완하기 위해 입력 임베딩에 위치 인코딩을 더하는 방식을 채택했다. 논문은 각 위치와 차원에 대해 사인·코사인 함수를 이용한 고정된 위치 인코딩을 기본으로 사용하며, 학습된 위치 임베딩과의 성능 차이가 거의 없음을 실험으로 확인했다.

"We also experimented with using learned positional embeddings [9] instead, and found that the two versions produced nearly identical results (see Table 3 row (E)). We chose the sinusoidal version because it may allow the model to extrapolate to sequence lengths longer than the ones encountered during training."
— 원문 링크

두 방식이 사실상 동등한 성능을 낸다면 왜 사인·코사인 방식을 택했을까? 저자들의 답은 '외삽 가능성'에 있다. 사인파 기반 인코딩은 훈련 중에 보지 못한, 더 긴 시퀀스에 대해서도 규칙에 따라 위치 값을 생성할 수 있기 때문이다. 이는 학습 데이터의 최대 길이를 넘어서는 입력에 대해 모델이 더 안정적으로 동작할 수 있는 근거를 제공한다.

"In Table 3 rows (B), we observe that reducing the attention key size dk hurts model quality. This suggests that determining compatibility is not easy and that a more sophisticated compatibility function than dot product may be beneficial. We further observe in rows (C) and (D) that, as expected, bigger models are better, and dropout is very helpful in avoiding over-fitting. In row (E) we replace our sinusoidal positional encoding with learned positional embeddings [9], and observe nearly identical results to the base model."
— 원문 링크

2-4. 기계 번역 벤치마크에서의 압도적 성능

WMT 2014 영어-독일어 번역 과제에서 트랜스포머 big 모델은 BLEU 28.4를 달성하며 앙상블을 포함한 기존 최고 모델보다 2 BLEU 이상 향상된 새로운 최고 성능을 수립했다. 더욱 주목할 만한 것은 효율성이다. 이 성과는 8개의 P100 GPU에서 3.5일 훈련으로 달성되었으며, 더 가벼운 base 모델은 동일 환경에서 단 12시간 만에 이전에 발표된 모든 모델과 앙상블을 능가했다.

"On the WMT 2014 English-to-German translation task, the big transformer model (Transformer (big) in Table 2) outperforms the best previously reported models (including ensembles) by more than 2.0 BLEU, establishing a new state-of-the-art BLEU score of 28.4. The configuration of this model is listed in the bottom line of Table 3. Training took 3.5 days on 8 P100 GPUs. Even our base model surpasses all previously published models and ensembles, at a fraction of the training cost of any of the competitive models."
— 원문 링크

WMT 2014 영어-프랑스어 번역에서도 big 모델은 BLEU 41.0을 기록하며 기존 단일 모델 최고 성능을 초과했다. 이때 소요된 훈련 비용은 이전 최고 모델의 4분의 1에도 미치지 않았다. 성능과 효율이라는 두 축 모두에서 트랜스포머의 우위가 입증된 셈이다.

"We trained our models on one machine with 8 NVIDIA P100 GPUs. For our base models using the hyperparameters described throughout the paper, each training step took about 0.4 seconds. We trained the base models for a total of 100,000 steps or 12 hours. For our big models,(described on the bottom line of table 3), step time was 1.0 seconds. The big models were trained for 300,000 steps (3.5 days)."
— 원문 링크

2-5. 아블레이션 연구: 아키텍처 선택의 근거

논문은 다양한 아블레이션 실험을 통해 각 설계 선택의 기여도를 검증했다. 어텐션 헤드 수와 키-값 차원을 변경한 실험에서, 단일 헤드 어텐션은 최적 설정 대비 BLEU가 0.9 낮았다. 반면 헤드 수를 과도하게 늘려도 성능이 저하되어 적절한 균형이 필요함을 보였다. 키 크기 dkd_k를 줄이면 모델 품질이 하락했는데, 이는 점곱 방식보다 더 정교한 호환성 함수가 유익할 수 있음을 시사한다.

"In Table 3 rows (A), we vary the number of attention heads and the attention key and value dimensions, keeping the amount of computation constant, as described in Section 3.2.2. While single-head attention is 0.9 BLEU worse than the best setting, quality also drops off with too many heads."
— 원문 링크

정규화 기법의 효과도 확인되었다. 레이블 스무딩(ϵ=0.1\epsilon=0.1)을 적용하면 모델 퍼플렉시티는 다소 높아지지만 정확도와 BLEU 점수는 개선된다. 모델이 더 '불확실하게' 학습함으로써 오히려 실제 평가 지표가 올라가는 이 역설적인 결과는, 정답 레이블에 지나치게 확신을 갖도록 학습시키는 것이 일반화에 해롭다는 점을 보여준다.

"In Table 3 rows (B), we observe that reducing the attention key size dk hurts model quality. This suggests that determining compatibility is not easy and that a more sophisticated compatibility function than dot product may be beneficial. We further observe in rows (C) and (D) that, as expected, bigger models are better, and dropout is very helpful in avoiding over-fitting. In row (E) we replace our sinusoidal positional encoding with learned positional embeddings [9], and observe nearly identical results to the base model."
— 원문 링크

2-6. 구문 파싱으로의 일반화: 번역을 넘어선 확장성

트랜스포머의 강점은 기계 번역에만 국한되지 않는다. 논문은 영어 구성 성분 파싱 과제를 통해 모델의 일반화 능력을 검증했다. 과제 특화 튜닝 없이도 순환 신경망 문법을 제외한 이전에 보고된 모든 모델을 능가하는 성능을 달성했다.

"Our results in Table 4 show that despite the lack of task-specific tuning our model performs surprisingly well, yielding better results than all previously reported models with the exception of the Recurrent Neural Network Grammar [8]."
— 원문 링크

특히 펜 트리뱅크의 훈련 데이터(약 4만 문장)만 사용한 소규모 실험에서도 트랜스포머는 버클리 파서를 능가했다. RNN 기반 시퀀스-투-시퀀스 모델이 동일 조건에서 이 성과를 달성하지 못했다는 사실과 대조하면, 트랜스포머의 범용 언어 구조 학습 능력이 단순한 데이터 규모의 산물이 아님을 알 수 있다. 어텐션 헤드들이 문장의 통사적·의미적 구조와 관련된 패턴을 스스로 학습한다는 해석 가능성 분석은 이를 정성적으로 뒷받침한다.

"In contrast to RNN sequence-to-sequence models [37], the Transformer outperforms the BerkeleyParser [29] even when training only on the WSJ training set of 40K sentences."
— 원문 링크

3. 시사점 및 한계

트랜스포머의 등장은 단순히 성능 지표 하나를 끌어올린 것이 아니라, 시퀀스 모델링의 패러다임을 근본적으로 바꾸었다. 순환 연산이 없으므로 전체 시퀀스를 병렬로 처리할 수 있고, 셀프 어텐션 덕분에 임의의 두 위치 사이를 상수 횟수의 연산으로 연결해 장거리 의존성을 효과적으로 포착한다. 가중치 공유, 사인파 위치 인코딩, 멀티 헤드 어텐션, 잔차 연결, 층 정규화 같은 설계 선택들은 이후 BERT, GPT 계열 모델이 공유하는 현대 언어 모델의 기본 문법이 되었다.

"To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution. In the following sections, we will describe the Transformer, motivate self-attention and discuss its advantages over models such as [17, 18] and [9]."
— 원문 링크

그러나 이 논문이 해결하지 못한 과제도 명확하다. 첫째, 셀프 어텐션의 연산 비용은 시퀀스 길이 nn에 대해 O(n2)O(n^2)로 증가한다. 긴 문서나 코드처럼 수천 토큰을 다루는 경우 이 이차 복잡도가 병목이 된다. 저자들도 매우 긴 시퀀스에 대해 지역화 전략을 적용하는 방향을 후속 연구로 남겨두었다. 둘째, 구성 성분 파싱 실험은 트랜스포머의 범용성을 보여주었지만, 당시 순환 신경망 문법이라는 구조화된 모델에는 미치지 못했다. 이는 명시적인 구조적 귀납 편향이 순수 어텐션 기반 모델에 비해 여전히 특정 과제에서 유리할 수 있음을 시사한다.

"As noted in Table 1, a self-attention layer connects all positions with a constant number of sequentially executed operations, whereas a recurrent layer requires O(n) sequential operations. In terms of computational complexity, self-attention layers are faster than recurrent layers when the sequence length n is smaller than the representation dimensionality d, which is most often the case with sentence representations used by state-of-the-art models in machine translations, such as word-piece [38] and byte-pair [31] representations. To improve computational performance for tasks involving very long sequences, self-attention could be restricted to considering only a neighborhood of size r in the input sequence centered around the respective output position. This would increase the maximum path length to O(n/r). We plan to investigate this approach further in future work."
— 원문 링크

이 논문의 진정한 유산은 '어텐션만으로 충분하다'는 증명이 열어준 가능성의 공간에 있다. 이후 등장한 수많은 대형 언어 모델들이 이 아키텍처를 기반으로 구축되었으며, 바이트 페어 인코딩이나 아담 옵티마이저 같은 학습 기법의 중요성도 이 논문의 실험적 검증을 통해 재확인되었다. 2017년의 이 선언은 현재까지도 AI 연구의 가장 중요한 전환점 중 하나로 남아 있다.

"We also experimented with using learned positional embeddings [9] instead, and found that the two versions produced nearly identical results (see Table 3 row (E)). We chose the sinusoidal version because it may allow the model to extrapolate to sequence lengths longer than the ones encountered during training."
— 원문 링크