HOW AI WORKS - FRAME BOOK
VOL. 01·SEP 2026

TRANSFORMERLAYER STACK.

A SEQUENCE OFTRANSFORMER LAYERSPROCESS INFORMATIONTHROUGH ATTENTION,FEED-FORWARD NETWORKS,AND RESIDUALCONNECTIONS.

h(l+1)=h(l)+f(h(l))h^{(l+1)} = h^{(l)} + f\left(h^{(l)}\right)
h(l)∈Rn×dh^{(l)} \in \mathbb{R}^{n \times d}
nn
= sequence length
dd
= model dimension
ll
= layer index
RESIDUAL STREAMh(L)h^{(L)}
h(0)h^{(0)}(residual stream input)

FORWARD

PASS

(inference)

h(l+1)h^{(l+1)}h(l)h^{(l)}Add & NormFeed-ForwardAdd & NormSelf-AttentionAdd & Norm

BACKWARD

PASS

(backpropagation)

∂L∂h(l+1)\dfrac{\partial L}{\partial h^{(l+1)}}∂L∂h(l)\dfrac{\partial L}{\partial h^{(l)}}Grad throughAdd & NormGrad throughFeed-ForwardGrad throughAdd & NormGrad throughSelf-AttentionGrad throughAdd & Norm

LAYER NORM

Normalize activations

for stable training.

FEED-FORWARD (MLP)

Non-linear transformation

applied token-wise.

SELF-ATTENTION

Global information flow

across sequence positions.

LAYER NORM

Pre-normalization

before attention.

  1. LAYER NORM

    Normalize activations for stable training.

  2. FEED-FORWARD (MLP)

    Non-linear transformation applied token-wise.

  3. SELF-ATTENTION

    Global information flow across sequence positions.

  4. LAYER NORM

    Pre-normalization before attention.

FORWARD PASS

(inference)

  1. h(l+1)h^{(l+1)}
  2. Add & Norm
  3. Feed-Forward
  4. Add & Norm
  5. Self-Attention
  6. Add & Norm
  7. h(l)h^{(l)}

BACKWARD PASS

(backpropagation)

  1. ∂L∂h(l+1)\dfrac{\partial L}{\partial h^{(l+1)}}
  2. Grad throughAdd & Norm
  3. Grad throughFeed-Forward
  4. Grad throughAdd & Norm
  5. Grad throughSelf-Attention
  6. Grad throughAdd & Norm
  7. ∂L∂h(l)\dfrac{\partial L}{\partial h^{(l)}}

ASELF-ATTENTION

Attention(Q,K,V)=\texttt{Attention}(Q, K, V) =
softmax(QKTdk)V\mathrm{softmax}\left(\dfrac{QK^{T}}{\sqrt{d_k}}\right) V

BFEED-FORWARD (MLP)

FFN(x)=W2 σ(W1x+b1)+b2\texttt{FFN}(x) = W_2\,\sigma(W_1 x + b_1) + b_2

(applied position-wise)

h(l)h^{(l)}F(h(l))F(h^{(l)})h(l+1)h^{(l+1)}

CRESIDUAL CONNECTION

h(l+1)=h(l)+F(h(l))h^{(l+1)} = h^{(l)} + F\left(h^{(l)}\right)

(skip connection)

DBACKPROPAGATION

∂L∂h(l)=∂L∂h(l+1)(I+∂F∂h(l))\dfrac{\partial L}{\partial h^{(l)}} = \dfrac{\partial L}{\partial h^{(l+1)}}\left(I + \dfrac{\partial F}{\partial h^{(l)}}\right)

Gradients flow in reverse

through the same structure.

TransformerLayers &Backpropagation

Motion Design Style-Frame Book

A VISUAL EXPLORATION OFTHE MATHEMATICS, MECHANISMS AND BEAUTYBEHIND MODERN ARTIFICIAL INTELLIGENCE.

FORWARD PASS

h(L)h^{(L)}h(0)h^{(0)}Residual streamLayer normFeed-forward (MLP)Self-attentionLayer norm

Sequence inputtokens (embeddings)

BACKWARD PASS

∂L/∂h(L)\partial L / \partial h^{(L)}
∂L∂h(0)\dfrac{\partial L}{\partial h^{(0)}}

Gradients toinput embeddings