HOW AI WORKS - FRAME BOOK
VOL. 01·SEP 2026

TERMSACROSSTHE BOOK.

  1. 01ATTENTION11
  2. 02REPRESENTATION19
  3. 03TRANSFORMATION20
  4. 04GENERATION27

A BACK COVER THAT DOUBLES AS AN INDEX: 77 TERMS, FOUR THEMES, EVERY PAGE ONE CLICK AWAY.

SYNTHESIS GLYPH

24 PAGES ON ONE RING. EACH CHORD JOINS TWO CHAPTERS THAT SHARE AN IDEA; PULSES RUN IN READING ORDER.

THE LANGUAGE-MODEL LOOP IN ONE LINE:

pθ(xt+1∣x≤t)p_\theta(x_{t+1} \mid x_{\le t})
=softmax⁡(WU ht(L))= \operatorname{softmax}\big(W_U\, h^{(L)}_t\big)

pages = 24terms = 77chords = 31loop = 8 s

HOVER OR FOCUS A PAGE NUMBERTO SEE ITS LINKS.

TAP A PAGE NUMBER TO OPEN IT.

softmax⁡ ⁣(QK⊤dk)V\operatorname{softmax}\!\left(\dfrac{QK^{\top}}{\sqrt{d_k}}\right)V

TOKENS SHARE CONTEXT BY WEIGHTED LOOKUP.

cos⁡(xi,xj)=xi⋅xj∥xi∥ ∥xj∥\cos(x_i, x_j) = \dfrac{x_i \cdot x_j}{\lVert x_i \rVert\, \lVert x_j \rVert}

MEANING BECOMES DISTANCE AND DIRECTION.

θ←θ−η ∇θL(θ)\theta \leftarrow \theta - \eta\, \nabla_{\theta} L(\theta)

LAYERS COMPUTE; GRADIENTS TEACH THEM.

xt+1∼softmax⁡(z/T)x_{t+1} \sim \operatorname{softmax}(z / T)

SAMPLE A STEP, APPEND, REPEAT.

Glossary

Terms Across the Book

77 KEY TERMS FROM PAGES 002 – 023, GROUPED BY THE BOOK'S FOUR THEMES. EACH ONE HAS A SHORT DEFINITION AND LINKS BACK TO THE PAGES WHERE THE IDEA IS DRAWN.

The book's four themes as a stack of layers — 04 generation: 27 terms, 03 transformation: 20 terms, 01 attention: 11 terms, 02 representation: 19 terms.

HOW TO READ

Embeddingx∈Rdx \in \mathbb{R}^{d}
A learned vector that stands for a token or an image patch.
Pages: 004017019010

THEMES

01 ATTENTION11
02 REPRESENTATION19
03 TRANSFORMATION20
04 GENERATION27
TOTAL77

GLOSSARY77 TERMS · 002 — 023

01ATTENTIONHOW TOKENS SHARE INFORMATION AND CONTEXT.11 TERMS

Attention
Each token builds its new vector as a weighted mix of other tokens' values.
Pages: 005001007
Query, key, valueQ, K, VQ,\ K,\ V
Three vectors per token: what it looks for, what it offers, and what it passes on.
Pages: 005014020
Attention head
One attention computation; several run in parallel, each learning its own pattern.
Pages: 005001
Causal mask
Hides later tokens, so each position can only attend to itself and earlier ones.
Pages: 014011
Positional encoding
A signal that tells the model where each token sits in the sequence.
Pages: 005019
Context window
The maximum number of tokens a model can take in at once.
Pages: 014
KV cache
Keys and values of earlier tokens, saved so each new token needs less computation.
Pages: 014
Grouped-query attention (GQA)
Several query heads share one set of keys and values, so the KV cache needs less memory.
Pages: 014
Prefill and decode
The two phases of inference: read the whole prompt in parallel, then write one token per step.
Pages: 014
Cross-attention
Attention where one sequence queries another, e.g. image features reading a text prompt.
Pages: 020
System prompt
Instructions placed before the conversation that set the model's role and rules.
Pages: 023014017

02REPRESENTATIONHOW TEXT, IMAGES AND MEANING BECOME VECTORS.19 TERMS

Token
A unit of text the model reads: a word, part of a word, or a symbol.
Pages: 004010
Tokenizer
Fixed rules that split text into tokens and give each one an ID.
Pages: 004010
Vocabulary∣V∣|V|
Every token a model knows, each with an integer ID (often 50,000 or more).
Pages: 004006011
Vector
An ordered list of numbers; also a point or a direction in space.
Pages: 002004
Embeddingx∈Rdx \in \mathbb{R}^{d}
A learned vector that stands for a token or an image patch.
Pages: 004017019010
Latent space
The space of a model's internal vectors, where related things lie close together.
Pages: 022001009020
Cosine similaritycos⁡(xi,xj)\cos(x_i, x_j)
How closely two vectors point the same way, from −1 to 1.
Pages: 017004
Retrieval-augmented generation (RAG)
Search a document store, often by embedding similarity, and add the best passages to the prompt.
Pages: 017021
Vector index
A store of embeddings built to find the ones nearest to a query, fast.
Pages: 017
Patch
A small square of an image, flattened and embedded like a token.
Pages: 019
Vision Transformer (ViT)
A transformer that reads an image as a sequence of patch embeddings, just as it reads text.
Pages: 019
CLIP
An image encoder and a text encoder trained so each picture lands near its caption in one shared space.
Pages: 019020
Residual stream
Each token's running vector; every layer reads it and adds its output back.
Pages: 007022
Featuredid_i
A direction in activation space that tracks a concept people can name.
Pages: 022
Polysemantic neuron
A single neuron that responds to several unrelated concepts.
Pages: 022
Superposition
Hypothesis that models store more features than dimensions as nearly orthogonal directions.
Pages: 022
Sparse autoencoder (SAE)
Rewrites an activation as a sum of a few features from a large learned dictionary.
Pages: 022
Probe
A small classifier trained on activations to test what information they contain.
Pages: 022
Steeringx+α dix + \alpha\, d_i
Adding a feature direction to activations to change what the model does.
Pages: 022

03TRANSFORMATIONHOW LAYERS COMPUTE AND HOW MODELS LEARN.20 TERMS

Neuron
A weighted sum of inputs plus a bias, passed through an activation function.
Pages: 006
Parametersθ\theta
The weights and biases a model learns; billions in a large model.
Pages: 006008012
Activation functionσ\sigma
A nonlinear function, such as ReLU, applied after a weighted sum.
Pages: 006
Transformer layer
Attention plus a feed-forward network, joined by residual connections and normalization.
Pages: 007010
Feed-forward network (MLP)
A small two-layer network applied to each token separately inside every layer.
Pages: 007
LossL(θ)L(\theta)
A number that measures how wrong the model is; training lowers it.
Pages: 008011012
Cross-entropy−log⁡p-\log p
The next-token loss: minus the log of the probability given to the true token.
Pages: 011021
Gradient∇L\nabla L
The direction in which the loss rises fastest: one slope per parameter.
Pages: 008007
Backpropagation
Computing every gradient layer by layer, backward, with the chain rule.
Pages: 007
Gradient descent
Repeatedly moving the parameters a small step against the gradient.
Pages: 008011
Learning rateη\eta
The size of each gradient-descent step.
Pages: 008
Pre-training
Training on a huge text corpus by predicting the next token.
Pages: 011012015
Scaling laws
Loss falls predictably, roughly as a power law, as parameters, data and compute grow.
Pages: 012
Compute-optimal training
For a fixed compute budget, grow parameters and data together: about 20 training tokens per parameter.
Pages: 012
Emergent abilities
Skills that seem to switch on suddenly at scale; partly an effect of all-or-nothing scoring.
Pages: 012
Fine-tuning
Further training of a pre-trained model on a smaller, targeted dataset.
Pages: 015
RLHF
Reinforcement learning from human feedback: training toward answers people prefer.
Pages: 015
Reward model
A model trained on people's choices between two answers, used to give any answer a score.
Pages: 015
KL penalty
A cost for drifting from the reference model, so RLHF gains reward without forgetting how to write.
Pages: 015
Alignment
Shaping a model's behavior to match human intentions and values.
Pages: 015023

04GENERATIONHOW OUTPUTS ARE PRODUCED, USED AND CHECKED.27 TERMS

Logitziz_i
The raw score a model gives each vocabulary token before softmax.
Pages: 006013
Softmaxezi/∑jezje^{z_i} / \textstyle\sum_j e^{z_j}
Turns a list of scores into positive probabilities that sum to 1.
Pages: 006013005
Next-token prediction
The core task: assign a probability to every possible next token.
Pages: 011006010
Autoregressive generation
Generate one token, append it to the input, repeat.
Pages: 010014013
Sampling
Picking the next token at random, weighted by its probability.
Pages: 013021016
Greedy decoding
Always picking the single most likely next token, so the same prompt gives the same text.
Pages: 013006
TemperatureTT
Divides the logits before softmax: low is focused, high is more varied.
Pages: 013021
Top-k
Keep only the k best candidates: the k likeliest tokens when sampling, or the k closest chunks in a search.
Pages: 013017
Top-p (nucleus) sampling
Sampling only from the smallest set of tokens whose probabilities add up to p.
Pages: 013
Chain of thought
Writing out intermediate reasoning steps as tokens before the answer.
Pages: 016
Test-time compute
Spending more computation while answering, by writing longer reasoning or sampling more attempts.
Pages: 016
Self-consistency
Sampling several chains of reasoning and keeping the answer most of them reach.
Pages: 016
Tool use
The model writes a structured call; software runs it and returns the result.
Pages: 018
Model Context Protocol (MCP)
An open standard for plugging tools and data sources into AI applications.
Pages: 018
Agent
A loop in which a model plans, acts with tools, observes the results, and repeats.
Pages: 018
Diffusion model
Generates an image by removing noise step by step, starting from pure noise.
Pages: 009020
Latent diffusion
Diffusion run in a compressed latent space instead of on pixels.
Pages: 020009
Classifier-free guidance
Pushes each denoising step toward the prompt by comparing predictions with and without it.
Pages: 020
Hallucination
Fluent, confident output that is false or not supported by any source.
Pages: 021017
Knowledge cutoff
The date the training data ends; the model has seen nothing that happened after it.
Pages: 021
Calibration
How well a model's confidence matches how often it is actually right.
Pages: 021
Benchmark
A fixed, held-out set of test cases with a scoring rule.
Pages: 023
pass@k
The chance that at least one of k sampled solutions passes the tests.
Pages: 023
Contamination
Test items leaking into training data, which inflates scores.
Pages: 023
LLM-as-a-judge
Using a model to compare or grade answers when there is no single correct one.
Pages: 023
Red-teaming
Deliberately attacking a model to find harmful or failing behavior.
Pages: 023
Safety classifier
A separate model that screens inputs or outputs and blocks unsafe ones.
Pages: 023

LATENT

How AI Works

A MOTION DESIGN STYLE-FRAME BOOK

COLOPHON

EDITION
VOL. 01 · SEP 2026
PAGES
001 — 024
CREDIT
Compiled & developed by Vu Dinh
TYPE
IBM Plex Mono · Inter Tight