Bldev's Blog

Interactive Architecture Lab

Hands-on empirical simulations of transformer bottlenecks, memory scaling, and hardware constraints

Computational Graph & Tensor Dataflow Engine

Computational Graph & Tensor Dataflow Engine

Attention Is All You Need (Vaswani et al., 2017)

Transformer Scaled Dot-Product Attention

Computes semantic token-to-token similarity (QKTQ K^T) and applies softmax weights over value representations

시퀀스 길이 (NN)512
64512 (Default)2048
헤드 차원 (dkd_k)64
3264 (Default)256
Target Mathematical FormulaStep 1 / 5
Attention(Q,K,V)=softmax(QKTdk)V\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{Q K^T}{\sqrt{d_k}}\right) V
Transpose Key Matrix

Transposes the [N×dkN \times d_k] Key matrix into [dk×Nd_k \times N] for dot-product multiplication

Current Sub-Operation
KT=Transpose(K)K^T = \text{Transpose}(K)
KK
[512×64][512 \times 64]
512 Rows × 64 Cols
transpose0 FLOPs (IO/Reshape)
KTK_T
[64×512][64 \times 512]
KTK^T [dk×Nd_k \times N]
Total Computation:68.16 MFLOPs
Output:[512×64][512 \times 64]

Core Architectural & Hardware Specification

Definition

A computational framework abstracting neural equations into a Directed Acyclic Graph (DAG), mapping tensors as dataflow edges and tracking autograd dependencies.

Core Mechanics
  • •Decomposes math into primitive nodes (MatMul, Softmax, Norm) and tensor edges
  • •Tracks backward paths by retaining forward activations to apply the chain rule
Limitations & Bottlenecks
  • •Dynamic control flows limit ahead-of-time static JIT optimizations
Comparison

Unlike text equations, computational graphs explicitly expose operator independence, enabling hardware compilers (XLA, TensorRT) to execute automatic parallelization.